Skip to content
TB
TeamBenchResources

How to Set Up Weighted Evaluation Criteria for Content

A practical guide to choosing criteria, assigning weights, and writing descriptions that produce useful AI review feedback — with examples for 6 content types.

TeamBench· Content Quality PlatformFebruary 10, 20269 min read

Weighted evaluation criteria are the foundation of any content review system. They define what "quality" means for your content, how important each dimension is, and what the AI reviewer should look for when scoring. Get the criteria right, and you get useful, actionable feedback. Get them wrong, and you get noise.

This guide covers how to choose criteria, assign weights, write effective descriptions, and configure criteria for different content types.

Quick answer: Choose 4-6 criteria based on the feedback you give most often. Assign weights that reflect real priorities (not equal weights). Write detailed descriptions that specify exactly what to evaluate. Test with 5-10 real pieces and adjust based on whether scores match your judgement.

How to Choose Your Criteria

Start With Your Feedback Patterns

The best criteria come from the feedback your editors actually give. Look at the last 20 pieces of content your team reviewed. What did editors comment on?

Common patterns:

  • "This doesn't sound like us" → Brand Voice criterion
  • "Too hard to read" / "Sentences are too long" → Readability criterion
  • "Where's the source for this claim?" → Accuracy criterion
  • "The keyword isn't in the title" → SEO Structure criterion
  • "What should the reader do next?" → CTA Effectiveness criterion
  • "This doesn't comply with the disclaimer requirement" → Compliance criterion

If editors comment on something in more than 30% of reviews, it's a criterion.

The Universal Five

Most content teams can start with these five criteria and customise from there:

  1. Brand Voice — does the content sound like your organisation?
  2. Readability — is the content clear and easy to consume?
  3. Accuracy — are facts, claims, and references correct?
  4. Structure — is the content well-organised for the format and channel?
  5. Effectiveness — does the content achieve its purpose (inform, persuade, convert)?

These five cover the dimensions that matter for almost every content type. The specifics — what "brand voice" means, what "readability" targets to use, what "structure" looks like — are defined in the criterion descriptions.

Industry-Specific Criteria

Some industries need additional or different criteria:

IndustryAdditional Criteria
HealthcareMedical accuracy, Patient privacy, Empathetic tone
Financial servicesRegulatory compliance, Risk language balance, Disclaimer presence
LegalJurisdictional accuracy, Privilege protection, Disclaimer language
EducationAccessibility, Learning objective alignment, Assessment validity
E-commerceFeature accuracy, Conversion optimisation, Product terminology

Download: Content Scoring Rubric Template

How to Assign Weights

The Priority Ranking Method

  1. List all your criteria
  2. Rank them from most important to least important for this content type
  3. Assign weights that roughly follow the ranking

Example — Blog post criteria:

RankCriterionWeight
1Brand Voice25%
2Readability25%
3Accuracy20%
4SEO Structure20%
5CTA Effectiveness10%
Total100%

The top criterion gets 25-30%. The bottom gets 10-15%. Everything else falls in between. The total must equal 100%.

Why Not Equal Weights?

Equal weights treat every criterion as equally important. That's rarely true. For a compliance-sensitive document, accuracy at 20% is underweighted — it should be 30-35%. For a social media post, SEO structure at 20% is overweighted — it should be 5-10% or removed entirely.

Equal weights also make overall scores less useful. If everything is weighted equally, a high brand voice score can mask a dangerously low accuracy score. Proper weighting ensures the overall score reflects what actually matters.

Content-Type-Specific Weights

Different content types need different weight distributions:

Blog Posts:

CriterionWeight
Brand Voice25%
Readability25%
Accuracy20%
SEO Structure20%
CTA Effectiveness10%

Marketing Emails:

CriterionWeight
Subject Line Quality25%
Clarity & Conciseness25%
CTA Effectiveness20%
Brand Voice20%
Personalisation10%

Product Descriptions:

CriterionWeight
Feature Accuracy30%
Benefit Focus25%
Brand Voice20%
SEO Optimisation15%
Scanability10%

Social Media Posts:

CriterionWeight
Brand Voice35%
Hook Quality25%
Clarity20%
CTA/Engagement20%

Compliance Documents:

CriterionWeight
Regulatory Accuracy35%
Disclaimer Completeness25%
Readability20%
Brand Voice10%
Formatting10%

Press Releases:

CriterionWeight
Newsworthiness25%
Accuracy25%
Structure (inverted pyramid)20%
Brand Voice15%
Quote Quality15%

How to Write Criterion Descriptions

The criterion description tells the AI reviewer exactly what to evaluate within each criterion. This is where most teams under-invest — and it's the #1 factor in feedback quality.

The Description Formula

Every criterion description should include:

  1. What to evaluate — the specific aspects of this dimension
  2. What good looks like — concrete targets and expectations
  3. What bad looks like — common failures to flag
  4. How to score — what merits a high score vs. a low score

Example: Brand Voice

Weak description:

"Check if the content matches our brand voice."

Strong description:

"Evaluate brand voice alignment across four dimensions:

  1. Tone: Our brand voice is confident, helpful, and human. Confident = we state things directly without hedging ('This reduces review time' not 'This might potentially help reduce review time'). Helpful = we explain why, not just what. Human = we write like a knowledgeable friend, not a corporate brochure.

  2. Terminology: Check against our preferred terms: 'content review' (not 'content audit'), 'criteria' (not 'parameters'), 'score' (not 'grade'), 'reviewer' (not 'checker'). Flag any use of banned terms: 'leverage', 'utilise', 'cutting-edge', 'industry-leading', 'best-in-class', 'in today's digital landscape'.

  3. Consistency: The tone should be consistent throughout. Flag any shifts — e.g., casual introduction then formal body, or confident opening then hedging conclusion.

  4. Voice markers: Active voice >80% of sentences. First person plural ('we') for company perspective. Second person ('you') when addressing the reader. Short sentences for impact. Questions to engage.

Score 90+: Perfect brand voice throughout, preferred terms used consistently, no banned terms, engaging personality evident. Score 70-89: Mostly aligned with minor deviations — 1-2 banned terms, occasional tone shifts, or slightly too formal. Score 50-69: Noticeable issues — multiple banned terms, inconsistent tone, reads more corporate than conversational. Score <50: Does not match brand voice — wrong tone, frequent banned terms, reads like generic AI output."

The strong description produces feedback that references specific paragraphs, flags specific terms, and suggests specific fixes. The weak description produces "Consider reviewing your brand voice" — useless.

Example: Readability

Strong description:

"Evaluate content readability targeting FK Grade 8-9 for blog posts:

  1. Sentence length: Average 15-20 words. Flag any sentence over 35 words. No more than 2 sentences over 25 words per paragraph.

  2. Paragraph length: 2-4 sentences per paragraph. Flag any paragraph over 5 sentences.

  3. Structure: Headings every 200-300 words. Clear H2/H3 hierarchy. Bulleted or numbered lists for sequences of 3+ items.

  4. Vocabulary: Prefer common words over technical jargon. Define any necessary technical terms on first use. Avoid nominalisations ('implementation' → 'implement', 'utilisation' → 'use').

  5. Active voice: >80% of sentences in active voice. Flag passive constructions where active would be clearer.

Score 90+: FK Grade 8-9, short clear sentences, excellent structure, no jargon. Score 70-89: FK Grade 9-11, mostly clear with a few long sentences or dense paragraphs. Score 50-69: FK Grade 11-13, multiple long sentences, insufficient headings, some jargon. Score <50: FK Grade 13+, very long sentences, wall-of-text paragraphs, heavy jargon."

Testing and Calibrating

The 10-Piece Test

Before deploying a reviewer, run 10 pieces of content through it:

  • 3 pieces you consider high quality (should score 80+)
  • 4 pieces of average quality (should score 65-75)
  • 3 pieces with known issues (should score below 65)

If scores don't match your expectations, adjust:

  • Scores too high across the board: Criterion descriptions are too lenient. Add more specific requirements and failure conditions.
  • Scores too low across the board: Criterion descriptions are too strict. Relax some requirements or lower the expectations for what constitutes a good score.
  • One criterion always scores the same: The description is too vague. Add more specific evaluation dimensions.

Calibration Adjustments

After the first month, review the data:

SignalAction
One criterion always scores 90+Description too lenient — tighten requirements
One criterion always scores <60Description too strict — or the team needs training
Two criteria have identical scoresThey're overlapping — merge or differentiate
Writers ignore feedback from one criterionThe feedback isn't actionable — rewrite the description
Overall scores don't match editor judgementWeight distribution needs adjustment

Getting Started

  1. List the top 5 things your editors check — these become your criteria
  2. Rank them by importance — assign weights accordingly
  3. Write a 4-5 sentence description for each — include what to evaluate, what good looks like, and how to score
  4. Create a reviewer with these criteria and a quality gate at 70
  5. Test with 10 pieces — calibrate based on results
  6. Deploy to the team — monitor score trends for the first month

Start now: Create a reviewer with custom criteria

Related:

evaluation-criteriaweighted-criteriacontent-scoringcontent-reviewrubriccontent-quality

Need consistent content quality across your team?

TeamBench lets you create custom AI reviewers that score content against your specific criteria. Submit content, get instant scored feedback, and improve with one click.

  • Create custom AI reviewers for your brand
  • Score content against your specific criteria
  • Instant feedback, one-click improvement
  • Free to start — no credit card required