Review Scores Seem Wrong
Quick Answer
Inaccurate scores are almost always caused by vague criteria guidance, missing knowledge base context, or misaligned weights. The fix is to make your criteria more specific, add relevant documentation, and calibrate against content you have already evaluated manually.
When review scores do not match your expectations, it is natural to question the AI's judgment. However, in the vast majority of cases, the issue is not the AI model but the instructions it is working from -- your criteria guidance text, weights, and available context.
The most common cause of inaccurate scores is vague criteria guidance. When guidance says "check for good quality," the AI has enormous latitude in interpretation, and its default interpretation may not match yours. Compare the feedback text with your expectations -- if the AI is evaluating something different from what you intended, the guidance needs to be more specific. Add concrete indicators, measurable targets, and examples of what constitutes high and low scores.
Missing knowledge base context is the second most common cause. If your reviewer does not have an attached knowledge base, the AI evaluates against general best practices rather than your specific standards. Content that is excellent by general standards might be off-brand by your standards. Adding your brand documentation to a knowledge base and attaching it to the reviewer often resolves the discrepancy immediately.
Misaligned weights can make the overall score feel wrong even when individual criterion scores are accurate. If you care deeply about accuracy but gave it a weight of 10% while formatting has 40%, a factually inaccurate but well-formatted piece will score surprisingly high overall. Review your weights and ensure they reflect your actual priorities.
Model selection matters for nuanced criteria. If you are using a basic model to evaluate subtle qualities like tone, voice, or persuasiveness, consider switching to a more capable model. Basic models excel at structural evaluations (readability, length, format) but may struggle with criteria that require deeper language understanding.
Check for the "anchor effect." If your content includes strong elements, the AI may give slightly higher scores across all criteria -- a well-structured piece might get a slightly generous brand voice score. This is a known tendency in AI evaluation and is mitigated by specific guidance text that forces the AI to focus on each criterion independently.
Run a calibration exercise. Take three pieces of content: one you consider excellent, one average, and one below your standards. Review all three with the same reviewer. The scores should clearly differentiate the three quality levels. If they do not, the issue is in the criteria, and you can pinpoint which criterion needs adjustment based on where the differentiation breaks down.
If you have tried all of the above and scores still feel off, contact support with specific examples of content and the scores you expected versus what you received. The team can help diagnose whether the issue is in the criteria configuration, the model selection, or a rare edge case that needs investigation.
Related Questions
Should I trust the overall score or individual criterion scores?
Individual criterion scores are more reliable indicators than the overall score. The overall score is a weighted average that can mask significant issues in one area. Always check per-criterion scores to understand where the content genuinely excels and where it needs work.
Do scores get more accurate over time?
The AI model does not learn from your specific reviews, but your scores get more accurate as you refine criteria guidance, add knowledge base content, and calibrate weights. The improvement comes from your configuration getting better.
Can I reset a reviewer and start fresh?
Yes, you can edit any reviewer at any time -- modify criteria, adjust weights, update guidance text, or replace it entirely. Previous reviews retain their original scores; changes only affect future reviews.
Related Free Tools
Still have questions?
Try TeamBench free and see how AI-powered content review works for your team.
Start Free Trial