Skip to content
TB
TeamBenchResources

How to Write AI Prompts for Content Review

Generic prompts give generic feedback. Here's how to structure AI review prompts with criteria, weights, and examples — plus why permanent reviewers beat one-off prompts.

TeamBench· Content Quality PlatformFebruary 9, 20269 min read

You paste your blog post into ChatGPT and ask "Is this good?" The response: "Yes, this is a well-written piece that covers the topic thoroughly. You might consider adding more examples in section 3." That's not useful feedback. It's AI politeness.

The problem isn't the AI — it's the prompt. A vague question gets a vague answer. A structured review prompt with specific criteria, weights, and examples gets specific, actionable feedback that actually improves your content.

Why Generic Prompts Fail

Generic PromptWhat You GetWhy It's Useless
"Is this good?""Yes, it's well-written."No specific feedback, no scoring, no priorities
"Review this blog post"A paragraph of general praise with 1-2 mild suggestionsAI defaults to positive — it's trained to be helpful, not critical
"Edit this for me"Minor word changes, comma adjustmentsSurface-level editing, misses structural and strategic issues
"How can I improve this?"5-10 suggestions of varying quality and relevanceNo prioritisation — which suggestion matters most?

The common thread: no criteria. Without defined standards, the AI has no framework for evaluation. It falls back on generic quality heuristics — which produce generic feedback.

The Structured Review Prompt

A review prompt that produces useful feedback has five components:

1. Role Definition

Tell the AI what kind of reviewer to be.

"You are a senior content editor reviewing blog posts for a B2B SaaS company. You are critical, specific, and focused on actionable improvements."

2. Criteria (The Most Important Part)

Define exactly what to evaluate. Each criterion should be specific enough that two reviewers would score the same content similarly.

"Evaluate this content against the following criteria:

  1. Argument clarity (weight: 3) — Is there a clear thesis? Does every paragraph advance the argument? Are claims connected logically?
  2. Evidence quality (weight: 3) — Are claims supported by specific data, examples, or sources? Count the number of unsupported claims.
  3. Brand voice (weight: 2) — Direct, knowledgeable, and approachable. No jargon without definition. No hedging.
  4. Readability (weight: 2) — Target Flesch-Kincaid grade 8-10. Flag sentences over 30 words and paragraphs over 4 sentences.
  5. SEO (weight: 2) — Primary keyword in title, H1, first 100 words, and at least one H2. Meta description present.
  6. CTA quality (weight: 1) — Clear, specific call to action that aligns with the content."

3. Scoring Framework

Tell the AI how to score.

"Score each criterion from 0-100. Calculate a weighted total score out of 100. For each criterion scoring below 70, provide 2-3 specific, actionable improvements with line references."

4. Context

Provide the context the AI needs to evaluate accurately.

"Target audience: Content managers at companies with 10-50 employees. Primary keyword: 'content review workflow'. The blog should position the reader's current process as inadequate and present a structured alternative."

5. Output Format

Specify exactly how you want the feedback structured.

"Format your response as:

  1. Overall score (weighted total)
  2. Score per criterion with brief justification
  3. Top 3 priority improvements (highest-weight criteria with lowest scores)
  4. Specific line-level feedback for the 5 weakest sections"

Complete Prompt Template

ROLE: You are a senior content editor. Be critical and specific.

CONTENT TO REVIEW:
[Paste your content here]

REVIEW CRITERIA:
1. [Criterion 1] (weight: [1-3]) — [What to evaluate]
2. [Criterion 2] (weight: [1-3]) — [What to evaluate]
3. [Criterion 3] (weight: [1-3]) — [What to evaluate]
4. [Criterion 4] (weight: [1-3]) — [What to evaluate]
5. [Criterion 5] (weight: [1-3]) — [What to evaluate]

CONTEXT:
- Target audience: [Who]
- Primary keyword: [Keyword]
- Content type: [Blog/email/case study/etc.]
- Brand voice: [Attributes]

SCORING:
Score each criterion 0-100. Calculate weighted total.
For any criterion below 70, provide specific improvements.

OUTPUT FORMAT:
1. Overall weighted score
2. Per-criterion scores with justification
3. Top 3 priority improvements
4. 5 specific line-level suggestions

Prompt Templates by Content Type

Blog Post Review Prompt

Review this blog post as a senior content editor.

CRITERIA:
1. Argument clarity (weight: 3) — Clear thesis, logical progression,
   every paragraph advances the argument
2. Evidence quality (weight: 3) — Claims supported by data/examples/sources,
   count unsupported assertions
3. Opening hook (weight: 2) — First paragraph earns continued reading
4. Brand voice (weight: 2) — [Your voice attributes]
5. Readability (weight: 2) — Grade 8-10, sentences under 25 words avg
6. SEO (weight: 2) — Keyword in title/H1/first 100 words/one H2
7. CTA (weight: 1) — Clear, specific next step for the reader

Score each 0-100. Weighted total. Top 3 improvements for lowest-scoring
high-weight criteria. 5 specific line-level fixes.

Email Copy Review Prompt

Review this marketing email as an email marketing specialist.

CRITERIA:
1. Subject line (weight: 3) — Under 50 chars, specific, benefit-clear
2. CTA clarity (weight: 3) — One CTA, action verb, low friction
3. Value delivery (weight: 2) — Opens with value, not self-promotion
4. Mobile readability (weight: 2) — Short paragraphs, scannable
5. Brand voice (weight: 2) — [Your voice attributes]
6. Personalisation (weight: 1) — Beyond just name — segment relevance

Score each 0-100. Flag the single biggest improvement opportunity.

Case Study Review Prompt

Review this case study as a B2B content strategist.

CRITERIA:
1. Customer centricity (weight: 3) — Story is about the customer,
   not about us
2. Data specificity (weight: 3) — Quantified results with before/after
3. Narrative quality (weight: 2) — Reads as a story with tension
   and resolution
4. Credibility (weight: 2) — Specific details that feel real,
   not generic
5. Brand voice (weight: 1) — [Your voice attributes]

Score each 0-100. Identify the weakest narrative section and suggest
a specific rewrite approach.

The Limitation of Prompt-Based Review

Prompt-based review works. But it has fundamental limitations compared to a permanent, configured reviewer:

FactorOne-Off PromptPermanent Reviewer
ConsistencyDifferent phrasing = different resultsSame criteria every time
ScoringAI interprets scoring differently each timeCalibrated scoring methodology
Quality gatesNo threshold enforcementPass/fail at defined thresholds
Trend trackingNo history — each review is isolatedScores tracked over time
Team standardisationEach person writes their own promptEveryone uses the same reviewer
Knowledge baseMust paste context every timePermanently uploaded brand guide, style guide
Time per review2-5 min to craft the prompt + paste content30 seconds to submit
Improvement over timeYou refine the prompt manuallyCriteria calibrated once, applied consistently

The progression: Start with prompts to learn what criteria matter. Then configure a permanent reviewer with those criteria for consistent, scalable, scored review.

Building Better Prompts: Common Mistakes

Mistake 1: No Criteria

"Review this content" without criteria is asking the AI to guess what you care about. It will default to grammar, clarity, and generic suggestions — missing what actually matters for your content.

Mistake 2: Too Many Criteria

More than 8 criteria dilutes focus. The AI provides shallow feedback across too many dimensions. Choose 5-8 criteria that matter most and weight them appropriately.

Mistake 3: No Weights

Without weights, the AI treats all criteria equally. A grammar error gets the same attention as a missing argument. Weights tell the AI where to focus.

Mistake 4: No Scoring

"Is this good?" produces a yes/no answer. "Score this 0-100 on each criterion" produces quantified feedback you can act on and track over time.

Mistake 5: No Examples

If you want the AI to evaluate brand voice, show it what your brand voice sounds like. Include Do/Don't examples in your prompt. Without examples, "approachable" means something different to every AI response.

Mistake 6: No Output Format

Without a specified format, the AI produces a free-form essay response. Specify: scores, prioritised improvements, and specific line-level fixes.

From Prompts to Permanent Reviewers

Once you've iterated on your prompts and know what criteria work:

Prompt ComponentBecomes
Role definitionSystem prompt for the reviewer
Criteria with weightsConfigured criteria in the reviewer
Context (brand guide, style guide)Knowledge base uploads
Scoring frameworkBuilt-in scoring methodology
Quality thresholdQuality gate configuration

The permanent reviewer does everything your prompt does — but consistently, at scale, with scoring history and quality gates.

Frequently Asked Questions

Can I use ChatGPT for content review?

Yes — with a structured prompt, ChatGPT provides useful feedback. The limitations: no consistent scoring, no quality gates, no trend tracking, and different responses for the same prompt. It works for ad-hoc review; it doesn't work for systematic quality assurance.

How long should a review prompt be?

200-400 words is ideal. Long enough to define criteria, context, and output format. Short enough that the AI focuses on reviewing your content rather than processing your instructions.

Should I include the full content in the prompt?

For content under 3,000 words, yes. For longer content, consider reviewing section by section — the AI maintains better focus on shorter sections.

How do I know if the AI's feedback is accurate?

Compare AI feedback to human editorial feedback on the same content. Where they agree, the criteria are well-defined. Where they diverge, either refine the criteria or note that specific dimension as requiring human judgment.

Can I use the same prompt for different content types?

Use the same structure but different criteria. A blog post review prompt and an email review prompt share the format (criteria, scores, output) but have different criteria appropriate to each content type.

Key Takeaways

  • Generic prompts produce generic feedback. Structured prompts with criteria, weights, and scoring produce actionable feedback.
  • Five components of a review prompt: role definition, criteria with weights, scoring framework, context, and output format.
  • Start with 5-8 weighted criteria — more dilutes focus, fewer misses important dimensions.
  • Include Do/Don't examples for subjective criteria like brand voice.
  • Prompt-based review has limitations: inconsistent scoring, no quality gates, no trend tracking, no team standardisation.
  • The progression: start with prompts to learn what criteria matter, then configure a permanent reviewer for consistent, scalable review.

This article is for informational purposes. AI capabilities evolve rapidly. Test review prompts with your specific content and refine based on the quality and accuracy of feedback received.

ai-promptscontent-reviewprompt-engineeringai-editingchatgpt-editingcontent-quality

Need consistent content quality across your team?

TeamBench lets you create custom AI reviewers that score content against your specific criteria. Submit content, get instant scored feedback, and improve with one click.

  • Create custom AI reviewers for your brand
  • Score content against your specific criteria
  • Instant feedback, one-click improvement
  • Free to start — no credit card required