Evidence literacy · Reviewed July 24, 2026
A positive study is not the same as a proven claim.
Learn to find the exact bridge between what researchers measured and what a marketer, clinic, creator, or headline says.
Step one
Turn the pitch into a testable sentence
“Supports recovery” is not a claim you can evaluate. Name the peptide, product, population, outcome, comparator, route, and time frame.
“CJC-1295 supports body composition.”
“In resistance-trained adults, a defined CJC-1295 formulation improves measured lean mass versus placebo after a specified period.”
Evidence ladder
Different studies answer different questions
The endpoint gap
Biomarkers can be real without proving the advertised result
CJC-1295 studies can support the statement that GH and IGF-1 increased under the study conditions. That does not, by itself, prove more muscle, less fat, better recovery, or longer life. Each of those is a separate outcome requiring its own evidence.
Registry literacy
A registered trial is a plan, not a result
ClinicalTrials.gov can show that a study exists, what it plans to measure, its status, sponsor, and whether results were submitted. Registration improves transparency. It does not mean the intervention worked, was safe, or completed as planned.
- Recruitment and completion status
- Primary outcomes and time frames
- Participant count and eligibility
- Results submitted versus no results posted
- Changes between original and current records
Seven-question check
Use this before repeating a peptide claim
- 01
What exact product, form, route, and population were studied?
- 02
Was the comparison randomized, blinded, and controlled?
- 03
Did the study measure a biomarker or something people actually feel or do?
- 04
How many participants completed the study, and for how long?
- 05
Was the outcome chosen in advance, and was the difference meaningful?
- 06
Do independent groups or larger studies find the same result?
- 07
Does the marketed claim stay inside the study’s real boundaries?
Ten papers can all be animal studies, share the same research group, or measure the same indirect marker. Read the design, not just the bibliography.
06 · Research question
Reconstruct the question the trial was designed to answer
A useful trial question names the population, intervention, comparator, outcome, and time horizon. Readers often call this PICO, with time added when needed. This structure prevents a result from floating away from its context. “Retatrutide reduced weight” is incomplete; the trial population, formulation, comparison, analysis, and week of measurement determine what the sentence means.
Check whether the paper’s primary question matches the promotional claim. A study designed to characterize pharmacokinetics may report exploratory weight changes, but it was not necessarily powered or controlled to establish a weight-management outcome. Secondary and exploratory results can generate valuable hypotheses while carrying greater risk of chance findings.
Also identify the unit of analysis. A trial may randomize individuals, clinics, body sites, or treatment periods. Treating multiple measurements from one participant as independent can exaggerate precision. The methods should make the design and analysis unit clear.
07 · Prospective plan
Use the protocol and registry to detect moving goalposts
A prospective protocol should define eligibility, intervention, comparator, outcomes, time points, sample-size assumptions, and analysis. Public registration adds a dated record that readers can compare with the final report. It helps reveal outcomes that disappeared, were added late, or changed priority after results were known.
ClinicalTrials.gov records evolve. Open the history and note the original submission, primary completion date, status changes, and results-posting dates. A completed status means data collection ended under the registry definition; it does not mean the result was positive or published. A study with no posted results still requires a publication or other reliable result source.
Changes are not automatically misconduct. Clinical research encounters legitimate operational and scientific issues. The question is whether changes were documented, justified, and made without exploiting knowledge of the results. Undisclosed switching from an unfavorable primary outcome to a favorable secondary outcome weakens confidence.
08 · Bias controls
Randomization, concealment, blinding, and comparison do different jobs
Randomization uses chance to assign participants and helps balance known and unknown prognostic factors. Allocation concealment prevents recruiters from predicting the next assignment. Blinding reduces differences in expectations, co-interventions, behavior, outcome assessment, and analysis. A paper can be randomized without adequately concealing allocation or fully blinding the people who influence the result.
The comparator determines the question. Placebo can estimate effects beyond expectation and study participation; usual care tests addition to current practice; an active comparator asks how interventions differ. A weak or inappropriate comparator can make a product look effective without showing that it improves on relevant care.
Some interventions cannot be fully blinded. In those cases, look for blinded outcome assessors, standardized co-interventions, objective outcomes where appropriate, and transparent discussion of bias. “Double blind” should not end the inquiry; the methods should say who was blinded and how masking was maintained.
09 · Population
Eligibility criteria define who the result can represent
Read age, diagnosis, disease severity, body-mass range, prior treatment, comorbidities, organ function, medication exclusions, and recruitment setting. A narrow population can produce a clean efficacy estimate while leaving uncertainty for people outside those criteria. Broad marketing often deletes these boundaries.
Baseline characteristics show the participants who actually entered, not just who was eligible. Compare groups for important imbalances and examine representation by sex, race, ethnicity, age, and disease history when relevant. Small subgroup counts cannot support confident conclusions simply because a table reports them.
Generalizability is not all or nothing. A trial can strongly establish efficacy in its enrolled population while providing a plausible but unproven expectation elsewhere. State the direct population first and label extrapolation as extrapolation.
10 · Endpoints
Primary, secondary, exploratory, surrogate, and clinical outcomes carry different weight
The primary endpoint usually drives sample size and the central statistical test. Secondary endpoints add information but may require adjustment for multiple comparisons. Exploratory outcomes can identify patterns for future study. A press release may lead with the largest favorable number even when it was not the primary question.
Clinical outcomes describe how people feel, function, survive, or use care. Biomarkers and imaging measures can be valuable and sometimes serve as validated surrogates. FDA notes that candidate surrogates remain under evaluation, while validated surrogates have stronger evidence that changing them predicts a specific clinical benefit. Validation is context-specific.
Use the measurement scale carefully. Ask whether it was validated, who assessed it, what direction is better, and what change is considered meaningful. A statistically detectable difference on an unfamiliar scale may not be noticeable to patients.
11 · Effect size
Read estimates and confidence intervals before p-values
An effect estimate tells the observed difference. The confidence interval describes statistical uncertainty under the model. A narrow interval can support a precise estimate; a wide interval may remain compatible with important benefit, little effect, or harm. The p-value addresses compatibility with a null hypothesis, not the magnitude, clinical importance, or absence of bias.
For continuous outcomes, examine mean change, between-group difference, units, baseline values, and variability. For events, compare absolute risks as well as relative measures. A 50% relative reduction can mean a fall from 2 in 100 to 1 in 100, an absolute difference of one percentage point. Both descriptions are accurate; the absolute difference often helps people understand practical impact.
Prespecified responder thresholds can complement averages, but dichotomizing a continuous measure loses information. Number needed to treat or harm can be useful when baseline risk and follow-up are comparable. Never detach a number from its time horizon.
12 · Analysis
Withdrawals and missing data can change the apparent answer
Participant flow shows who started, stopped treatment, withdrew, was lost, or remained in the analysis. If adverse effects or lack of benefit cause more departures in one group, analyzing only completers can make the intervention look better. The reason for missingness matters as much as the percentage.
Trials use different estimands and methods to define what treatment effect they are estimating. One analysis may estimate the effect regardless of discontinuation or rescue therapy; another may focus on outcomes while participants remained on treatment. Both can be informative, but they answer different questions and can produce different numbers.
Look for sensitivity analyses that test alternative missing-data assumptions. A result that changes materially under plausible assumptions is less robust. Press releases often present one estimand without enough context, which is why full reports remain important.
13 · Safety
Count exposure, events, discontinuations, and observation time
Safety tables should be read by group and event severity. Common adverse events, serious events, events leading to discontinuation, laboratory changes, vital signs, and deaths each add information. An event after treatment is not automatically caused by treatment, but imbalances and patterns require assessment.
Sample size and duration set the detection limit. If an event occurs naturally once in several thousand exposures, a trial with a few hundred participants may miss it entirely. Long latency effects may not appear during short follow-up. The absence of a signal is reassuring only within the amount and type of observation available.
Trial safety also assumes the studied product and monitoring. It cannot establish the purity or sterility of an unrelated item, and it may not generalize to excluded populations or combinations. Keep clinical safety evidence and marketplace product risk as separate dimensions.
14 · Reporting
Peer review, sponsorship, replication, and the total record complete the appraisal
CONSORT 2025 provides a 30-item framework for transparent randomized-trial reporting, including participant flow. Reporting quality allows appraisal; it does not guarantee flawless design. A well-written paper can still have a weak comparator, short duration, or narrow population.
Funding and author relationships should be recorded, not used as automatic disqualifiers. Check who designed the study, controlled data, performed analysis, and decided to publish. Independent replication increases confidence that a finding is not specific to one sponsor, team, protocol, or dataset.
Finally, search beyond the paper. Regulatory reviews, registry results, later trials, corrections, and systematic syntheses can change the picture. A responsible conclusion describes where the study sits in the total evidence and what remains unknown.
15 · Worked interpretation
Turn a dramatic headline back into a study result
Suppose a headline says an investigational peptide produced “more than 20% weight loss.” Find the trial and identify the analysis population, maintenance group, comparator, time point, and estimand. Determine whether the number is a raw mean, model-based estimate, treatment-regimen result, or effect regardless of discontinuation. Check the placebo change and confidence interval rather than repeating the treatment number alone.
Then inspect participant flow and safety. How many people stopped treatment, and why? Were gastrointestinal events common? Did the report include serious events and changes in heart rate or laboratory measures? A strong average efficacy result and meaningful treatment burden can coexist. The article should communicate both.
Finally, classify the source. A peer-reviewed full report permits deeper appraisal than a topline sponsor release. A registry record can confirm the protocol and status but not the outcome. The correct public sentence should preserve those layers: who reported the result, which participants and product were studied, what the comparison showed, and which details remain unavailable.
This reconstruction is slower than sharing a headline and faster than correcting years of misinformation. Once practiced, it becomes a compact routine: question, protocol, population, endpoint, estimate, uncertainty, attrition, safety, source type, and boundary.
Primary references