The rating fidelity problem: how we guarantee AI never alters a condition rating
A RICS condition rating carries professional and legal weight. Otto's RatingFidelityChecker ensures the AI can never change a rating set by the surveyor. Here is how it works.
The rating fidelity problem: how we guarantee AI never alters a condition rating
When we were designing Otto, one question kept coming back: how do we make it impossible for the AI to change a condition rating that the surveyor set?
This is not an edge case. It is the central safety question for any AI tool that writes RICS reports, and one we had to answer before Otto could go anywhere near a real inspection (see why we built Otto in the first place). We built a component called RatingFidelityChecker specifically to answer it, and this post explains why we built it and how it works.
Why condition ratings are the load-bearing part
A RICS condition rating is not just a summary label. In a Level 2 HomeBuyer Report or a Level 3 Building Survey, each element of the property is assigned one of three condition ratings:
- Condition 1: No repair needed at present
- Condition 2: Defects requiring repair or replacement but not considered urgent
- Condition 3: Defects that are serious or require urgent repair
These ratings are what buyers act on. A Rating 3 on the roof structure can halt a transaction, prompt a price renegotiation, or lead to further specialist investigation. They are the professional judgement of a chartered surveyor, and they carry weight in any dispute or complaint.
If an AI model silently altered a Rating 3 to a Rating 2 while rephrasing a paragraph, the published report would not reflect the surveyor's findings. That is a professional liability risk and a potential harm to the buyer. We did not want to get anywhere near it.
What RatingFidelityChecker does
The fidelity checker runs as a verification pass between the AI-generated draft and the final document presented to the surveyor for sign-off.
It works in two directions. When the surveyor sets a condition rating (whether directly, or by confirming a suggestion from Otto), that rating is recorded as the authoritative value for that element. When the AI generates or regenerates text for that section, the checker compares every condition rating referenced in the output against the authoritative record.
If the AI output contains a rating that differs from what the surveyor set, the checker flags it. The discrepancy is surfaced to the surveyor for review rather than silently passed through. The surveyor decides whether the AI's wording was a genuine error, a transcription slip, or their own change of mind following review.
Nothing is overridden automatically in either direction. The checker's role is to catch disagreements, not to resolve them.
Why we did not rely on prompting alone
An earlier approach was to instruct the AI model in its prompt never to alter condition ratings. That instruction works most of the time. But language models are probabilistic, and the word "most" is not good enough when the output is a professional document with legal standing.
A prompt is a strong nudge. RatingFidelityChecker is a hard check. The distinction matters.
We treat the fidelity check the same way we think about validation in any other part of a system: a well-written instruction is useful, but it is not a substitute for a test that verifies the output. It is the same principle behind how we stopped Otto's sibling tool from making up property prices, a hard check against a source of truth, not a hope that the model behaves.
What this means for surveyors in practice
In day-to-day use, the fidelity checker is invisible. If the AI draft matches the ratings the surveyor set, the check passes silently and the surveyor reviews the report as normal.
It becomes visible only when there is a discrepancy, which in practice is uncommon but not impossible. The surveyor sees exactly where the mismatch is, what the AI wrote and what they had set, and they resolve it.
The design intent is simple: the surveyor's condition ratings are theirs. Otto helps them write; it does not judge. That applies whether a surveyor is producing a Level 2 HomeBuyer Report or a fuller Level 3 Building Survey: the ratings never move without the surveyor's say-so.
More about how Otto is built at RICS surveyors on Home.
Otto is a drafting tool. The chartered surveyor is responsible for all condition ratings and conclusions in any RICS report completed using Otto.
Thinking about your next move?
See what your home's worth in seconds, with the UK's longest-running property data behind it.
Get an instant valuation