How to Write AI Prompts for Content Review
Generic prompts give generic feedback. Here's how to structure AI review prompts with criteria, weights, and examples — plus why permanent reviewers beat one-off prompts.
You paste your blog post into ChatGPT and ask "Is this good?" The response: "Yes, this is a well-written piece that covers the topic thoroughly. You might consider adding more examples in section 3." That's not useful feedback. It's AI politeness.
The problem isn't the AI — it's the prompt. A vague question gets a vague answer. A structured review prompt with specific criteria, weights, and examples gets specific, actionable feedback that actually improves your content.
Why Generic Prompts Fail
| Generic Prompt | What You Get | Why It's Useless |
|---|---|---|
| "Is this good?" | "Yes, it's well-written." | No specific feedback, no scoring, no priorities |
| "Review this blog post" | A paragraph of general praise with 1-2 mild suggestions | AI defaults to positive — it's trained to be helpful, not critical |
| "Edit this for me" | Minor word changes, comma adjustments | Surface-level editing, misses structural and strategic issues |
| "How can I improve this?" | 5-10 suggestions of varying quality and relevance | No prioritisation — which suggestion matters most? |
The common thread: no criteria. Without defined standards, the AI has no framework for evaluation. It falls back on generic quality heuristics — which produce generic feedback.
The Structured Review Prompt
A review prompt that produces useful feedback has five components:
1. Role Definition
Tell the AI what kind of reviewer to be.
"You are a senior content editor reviewing blog posts for a B2B SaaS company. You are critical, specific, and focused on actionable improvements."
2. Criteria (The Most Important Part)
Define exactly what to evaluate. Each criterion should be specific enough that two reviewers would score the same content similarly.
"Evaluate this content against the following criteria:
- Argument clarity (weight: 3) — Is there a clear thesis? Does every paragraph advance the argument? Are claims connected logically?
- Evidence quality (weight: 3) — Are claims supported by specific data, examples, or sources? Count the number of unsupported claims.
- Brand voice (weight: 2) — Direct, knowledgeable, and approachable. No jargon without definition. No hedging.
- Readability (weight: 2) — Target Flesch-Kincaid grade 8-10. Flag sentences over 30 words and paragraphs over 4 sentences.
- SEO (weight: 2) — Primary keyword in title, H1, first 100 words, and at least one H2. Meta description present.
- CTA quality (weight: 1) — Clear, specific call to action that aligns with the content."
3. Scoring Framework
Tell the AI how to score.
"Score each criterion from 0-100. Calculate a weighted total score out of 100. For each criterion scoring below 70, provide 2-3 specific, actionable improvements with line references."
4. Context
Provide the context the AI needs to evaluate accurately.
"Target audience: Content managers at companies with 10-50 employees. Primary keyword: 'content review workflow'. The blog should position the reader's current process as inadequate and present a structured alternative."
5. Output Format
Specify exactly how you want the feedback structured.
"Format your response as:
- Overall score (weighted total)
- Score per criterion with brief justification
- Top 3 priority improvements (highest-weight criteria with lowest scores)
- Specific line-level feedback for the 5 weakest sections"
Complete Prompt Template
ROLE: You are a senior content editor. Be critical and specific.
CONTENT TO REVIEW:
[Paste your content here]
REVIEW CRITERIA:
1. [Criterion 1] (weight: [1-3]) — [What to evaluate]
2. [Criterion 2] (weight: [1-3]) — [What to evaluate]
3. [Criterion 3] (weight: [1-3]) — [What to evaluate]
4. [Criterion 4] (weight: [1-3]) — [What to evaluate]
5. [Criterion 5] (weight: [1-3]) — [What to evaluate]
CONTEXT:
- Target audience: [Who]
- Primary keyword: [Keyword]
- Content type: [Blog/email/case study/etc.]
- Brand voice: [Attributes]
SCORING:
Score each criterion 0-100. Calculate weighted total.
For any criterion below 70, provide specific improvements.
OUTPUT FORMAT:
1. Overall weighted score
2. Per-criterion scores with justification
3. Top 3 priority improvements
4. 5 specific line-level suggestions
Prompt Templates by Content Type
Blog Post Review Prompt
Review this blog post as a senior content editor.
CRITERIA:
1. Argument clarity (weight: 3) — Clear thesis, logical progression,
every paragraph advances the argument
2. Evidence quality (weight: 3) — Claims supported by data/examples/sources,
count unsupported assertions
3. Opening hook (weight: 2) — First paragraph earns continued reading
4. Brand voice (weight: 2) — [Your voice attributes]
5. Readability (weight: 2) — Grade 8-10, sentences under 25 words avg
6. SEO (weight: 2) — Keyword in title/H1/first 100 words/one H2
7. CTA (weight: 1) — Clear, specific next step for the reader
Score each 0-100. Weighted total. Top 3 improvements for lowest-scoring
high-weight criteria. 5 specific line-level fixes.
Email Copy Review Prompt
Review this marketing email as an email marketing specialist.
CRITERIA:
1. Subject line (weight: 3) — Under 50 chars, specific, benefit-clear
2. CTA clarity (weight: 3) — One CTA, action verb, low friction
3. Value delivery (weight: 2) — Opens with value, not self-promotion
4. Mobile readability (weight: 2) — Short paragraphs, scannable
5. Brand voice (weight: 2) — [Your voice attributes]
6. Personalisation (weight: 1) — Beyond just name — segment relevance
Score each 0-100. Flag the single biggest improvement opportunity.
Case Study Review Prompt
Review this case study as a B2B content strategist.
CRITERIA:
1. Customer centricity (weight: 3) — Story is about the customer,
not about us
2. Data specificity (weight: 3) — Quantified results with before/after
3. Narrative quality (weight: 2) — Reads as a story with tension
and resolution
4. Credibility (weight: 2) — Specific details that feel real,
not generic
5. Brand voice (weight: 1) — [Your voice attributes]
Score each 0-100. Identify the weakest narrative section and suggest
a specific rewrite approach.
The Limitation of Prompt-Based Review
Prompt-based review works. But it has fundamental limitations compared to a permanent, configured reviewer:
| Factor | One-Off Prompt | Permanent Reviewer |
|---|---|---|
| Consistency | Different phrasing = different results | Same criteria every time |
| Scoring | AI interprets scoring differently each time | Calibrated scoring methodology |
| Quality gates | No threshold enforcement | Pass/fail at defined thresholds |
| Trend tracking | No history — each review is isolated | Scores tracked over time |
| Team standardisation | Each person writes their own prompt | Everyone uses the same reviewer |
| Knowledge base | Must paste context every time | Permanently uploaded brand guide, style guide |
| Time per review | 2-5 min to craft the prompt + paste content | 30 seconds to submit |
| Improvement over time | You refine the prompt manually | Criteria calibrated once, applied consistently |
The progression: Start with prompts to learn what criteria matter. Then configure a permanent reviewer with those criteria for consistent, scalable, scored review.
Building Better Prompts: Common Mistakes
Mistake 1: No Criteria
"Review this content" without criteria is asking the AI to guess what you care about. It will default to grammar, clarity, and generic suggestions — missing what actually matters for your content.
Mistake 2: Too Many Criteria
More than 8 criteria dilutes focus. The AI provides shallow feedback across too many dimensions. Choose 5-8 criteria that matter most and weight them appropriately.
Mistake 3: No Weights
Without weights, the AI treats all criteria equally. A grammar error gets the same attention as a missing argument. Weights tell the AI where to focus.
Mistake 4: No Scoring
"Is this good?" produces a yes/no answer. "Score this 0-100 on each criterion" produces quantified feedback you can act on and track over time.
Mistake 5: No Examples
If you want the AI to evaluate brand voice, show it what your brand voice sounds like. Include Do/Don't examples in your prompt. Without examples, "approachable" means something different to every AI response.
Mistake 6: No Output Format
Without a specified format, the AI produces a free-form essay response. Specify: scores, prioritised improvements, and specific line-level fixes.
From Prompts to Permanent Reviewers
Once you've iterated on your prompts and know what criteria work:
| Prompt Component | Becomes |
|---|---|
| Role definition | System prompt for the reviewer |
| Criteria with weights | Configured criteria in the reviewer |
| Context (brand guide, style guide) | Knowledge base uploads |
| Scoring framework | Built-in scoring methodology |
| Quality threshold | Quality gate configuration |
The permanent reviewer does everything your prompt does — but consistently, at scale, with scoring history and quality gates.
Frequently Asked Questions
Can I use ChatGPT for content review?
Yes — with a structured prompt, ChatGPT provides useful feedback. The limitations: no consistent scoring, no quality gates, no trend tracking, and different responses for the same prompt. It works for ad-hoc review; it doesn't work for systematic quality assurance.
How long should a review prompt be?
200-400 words is ideal. Long enough to define criteria, context, and output format. Short enough that the AI focuses on reviewing your content rather than processing your instructions.
Should I include the full content in the prompt?
For content under 3,000 words, yes. For longer content, consider reviewing section by section — the AI maintains better focus on shorter sections.
How do I know if the AI's feedback is accurate?
Compare AI feedback to human editorial feedback on the same content. Where they agree, the criteria are well-defined. Where they diverge, either refine the criteria or note that specific dimension as requiring human judgment.
Can I use the same prompt for different content types?
Use the same structure but different criteria. A blog post review prompt and an email review prompt share the format (criteria, scores, output) but have different criteria appropriate to each content type.
Key Takeaways
- Generic prompts produce generic feedback. Structured prompts with criteria, weights, and scoring produce actionable feedback.
- Five components of a review prompt: role definition, criteria with weights, scoring framework, context, and output format.
- Start with 5-8 weighted criteria — more dilutes focus, fewer misses important dimensions.
- Include Do/Don't examples for subjective criteria like brand voice.
- Prompt-based review has limitations: inconsistent scoring, no quality gates, no trend tracking, no team standardisation.
- The progression: start with prompts to learn what criteria matter, then configure a permanent reviewer for consistent, scalable review.
This article is for informational purposes. AI capabilities evolve rapidly. Test review prompts with your specific content and refine based on the quality and accuracy of feedback received.