
How to Humanize a Systematic Review Draft After ChatGPT
ChatGPT systematic reviews invent studies and blur inclusion criteria. Lock the protocol, screened set, and extraction table first. Humanize synthesis voice once–never invent eligible papers.
A systematic review is not a literature review with a longer bibliography. It is a protocol-bound map of studies that met inclusion criteria you defined in advance. ChatGPT will happily “include” (Chen et al., 2018) for a trial that never existed and smooth every result into “significant improvements across outcomes.”
Humanizing a systematic review means polishing synthesis and discussion voice after the screened set and extraction table are locked. It does not mean rewriting Methods until PRISMA numbers drift.
Hubs: ethical AI use checklist, literature review humanize, citations, research papers.
What reviewers punish
| Element | Weak AI draft | Strong systematic review |
|---|---|---|
| Protocol | Vague “we searched widely” | Databases, dates, registers named |
| Inclusion | Flexible after results | Locked criteria; deviations reported |
| PRISMA counts | Round numbers that “sound right” | Match screening log exactly |
| Extraction | Narrative only | Table from papers you opened |
| Synthesis | Every study “important” | Heterogeneity and disagreement named |
| Citations | Plausible ghosts | Every row openable |
Detectors may flag smooth Discussion prose. Examiners flag fake included studies first.
Lock inclusion criteria before any model
Write the criteria by hand (or paste from your registered protocol):
| Domain | Your locked rule (example) |
|---|---|
| Population | Adults with X diagnosis |
| Intervention / exposure | Y as defined |
| Comparator | Usual care or Z |
| Outcomes | Primary outcome named |
| Study designs | RCTs only (example) |
| Languages / dates | As in protocol |
| Exclusions | Case reports, conference abstracts only–if protocol says so |
Do not invent studies to fill a thin yield. Empty cells and honest “no eligible studies for subgroup Q” beat hallucinated rows.
Build the extraction table only from PDFs you screened in:
| Study ID | Design | n | Intervention | Primary result | Risk of bias notes | PDF open? |
Prompt with the table and: “Draft Results synthesis from these rows only. Do not add studies, years, or effect sizes. Name where results conflict.”
PRISMA and Methods: what not to humanize away
| Section | Humanize? | Why |
|---|---|---|
| Database names, search dates | No | Audit trail |
| n screened / included / excluded | No | Must match log |
| Inclusion / exclusion lists | No (wording light only) | Protocol fidelity |
| Extraction numbers | No | From software / hand extract |
| Narrative synthesis paragraphs | Yes, one pass | Where AI flattens debate |
| Discussion limitations | Yes, after true limits listed | Cut “more research needed” mush |
| Reference list | No | Accuracy |
Paste stiff synthesis into Human Writes once. Put back study IDs, years, and statistics exactly. Re-open every cited PDF.
Before and after
AI Results sludge:
Numerous high-quality studies have demonstrated significant benefits of the intervention. Smith (2019) and hypothetical cohorts alike highlight promising outcomes, underscoring the importance of the approach across settings.
After extraction table + one pass:
Three RCTs met inclusion criteria (n = 412 total). Ahmad (2020) and Li (2021) found modest gains on the primary score at 12 weeks; Ortiz (2019) found no difference once baseline severity was adjusted. Between-study heterogeneity was high (I² reported from your software)–pooling was descriptive only.
Same assignment type. Second version survives “show me the included list.”
Workflow
- Freeze protocol / registration details before drafting prose.
- Complete screening log; export PRISMA counts from your tool.
- Fill extraction table from open PDFs only.
- Draft Results from the table–ban new citations in the prompt.
- Write Discussion limits you actually observed (search gaps, bias, sparse n).
- One Human Writes pass on synthesis and Discussion framing.
- Diff-check every number and study ID against the table.
- Supervisor review; disclose AI if policy requires (ethical checklist).
Related methods writing: lab report, research abstract, annotated bibliography.
What not to do
- Let ChatGPT “find additional key papers” for the included set.
- Humanize until effect sizes or confidence intervals change.
- Hide empty subgroups behind vague “the literature shows.”
- Copy PRISMA flowchart numbers from a template blog post.
- Treat systematic review voice like a persuasive essay (argumentative essay is a different genre).
Discussion without inventing implications
AI Discussion sections love “these findings have important implications for policy, practice, and future research” with no tether to your included set. Anchor Discussion to:
- What the included studies actually measured
- Where they disagree (population, dose, follow-up)
- Limits from your search and bias assessment
- One cautious implication you can defend in viva / defense
| Discussion move | Allowed | Not allowed |
|---|---|---|
| Compare two included trials | Yes | Cite a third trial you excluded without saying so |
| Note sparse evidence | Yes | Invent a “growing consensus” |
| Suggest research gap | Yes | Promise clinical guidelines your review did not support |
Humanize the sentences that sound like every other review’s closing page–not the claims underneath. Soft polish: Human Writes after the gap list is handwritten.
For narrative lit reviews without a protocol, use literature review–different genre, same ban on ghost citations. Broader ethics: ethical AI checklist.
Supervisor red flags to self-check
- Included n does not match the table
- Search date in prose ≠ Methods
- Risk-of-bias tool named but never applied in text
- Forest plot numbers that do not appear in extraction
- Perfect parallel “Author (Year) found that…” stacks with no debate
Fix facts before any second voice pass. A smoother wrong review is still wrong.
Bottom line
Systematic reviews win on protocol fidelity and a true included set. Human Writes is the voice pass after inclusion criteria and the extraction table are locked–never a license to invent studies.
Paste stiff synthesis on Human Writes when every cited paper is one you screened and can open.