EngineeringWhen Output Checks Threw Away Good AI Answers
· FINCH
- #llm
- #guardrails
- #prompt-engineering
Before FINCH's AI server sends an explanation to the user, the text goes through 10 automatic checks. They look at things like whether the model wrote a number directly, whether each footnote points to a real source, whether there's trading advice such as "a good time to buy", and whether the sentence count is in range. If any check fails, the server regenerates with the reasons attached, and if that fails too, the section isn't sent.
Strict checks stop wrong explanations, but they also throw away good ones, and every discarded section costs more tokens to regenerate. Once we were in production and counted the rejection reasons, a good share turned out to be false positives: text that was fine but got flagged anyway. This post is about finding and reducing them.
Part of the top rejection was a footnote format issue
Early in production, the top rejection reason was "raw number" (a number written directly instead of through a placeholder) with 137 cases, and second was the citation check with 102. Following up on an issue the backend part raised, it turned out both had the same root.
Footnotes are supposed to look like [^cit_5], but the model sometimes wrote ^cit_5 without the brackets. The number check removes footnotes first and then looks for digits. A bracketless footnote wasn't removed, so the digit 5 was left behind and flagged as a raw number.
Instead of asking the model to follow the format more carefully, the server now normalizes footnotes in the text to the standard form. The check actually got stricter: before, a bracketless footnote was invisible to it, so a made-up citation could slip through. Now it's caught.
When the prompt and the check disagree
Raw-number rejections kept coming after the footnote fix. Counting regeneration reasons over 6 hours of production, most of the 21 raw-number rejections were false positives.
12
'1' as a rank or unit number
6
'2026' as a date or year
3
', 주' and '한 주' (one week)
The prompt told the model to write years as-is, but the check only allowed the form "2026년". The quantity regex could start with a comma, so it read ", 주" as a quantity. Dates, years, ranks, equipment numbers and "한 주" (one week, as a period) went into the allowlist, and the regex now has to start with a digit (\d[\d,]*).
The prompt also got a section called "expressions that get discarded by automatic checks", listing the exact words the checker looks for. If the model doesn't know what gets thrown away, it keeps making the same mistake.
A fix that overcorrected
Sentence counts went the same way. An earlier prompt line, "it's safer to keep it short", led to overcorrection. In the batch run with the new prompt, 26 of 145 sections were discarded, and the top reason was stock analysis failing "at least 3 sentences required: 1". Portfolio diagnosis went the other way, running 5 to 9 sentences. Count words like "한 종목" (one stock) and "네 가지" (four things) were also rejected as quantities.
So the prompt now states the sentence counts exactly as the check enforces them. Stock analysis: "exactly 3 sentences, 4 only when there's plenty of evidence; 1–2 or 5 gets discarded". Diagnosis: exactly 3 sentences per finding. It also says that too few and too many are both discarded, and that count words are quantities too.
Three sentences counted as one
In the next batch, 13 of 35 discards were still "at least 3 sentences required: 1". This time it wasn't the model. It had written 3 sentences but put the footnote after the period.
…늘었습니다.[^cit_1] 다음 문장은…The sentence splitter only treated "terminal punctuation followed by whitespace" as an end, so a period followed by a footnote didn't count. The whole paragraph was counted as one sentence. Now footnotes after a period are moved in front of it before counting.
_SENTENCE_END_RE = re.compile(r"(?<!\d)[.!?]+(?=\s|$)")
_TRAILING_CITATION_RE = re.compile(r"([.!?]+)((?:\s*\[\^[^\]]+\])+)")
def split_sentences(text: str) -> list[str]:
text = _TRAILING_CITATION_RE.sub(lambda m: m.group(2).strip() + m.group(1), text)
...This fix changed only the check, not the prompt. A prompt change bumps the prompt version and invalidates every cached analysis. Changing only the check kept the cache, and the next batch rebuilt just the sections that had failed.
False positives in chat
The same day, after the backend relay was connected, the very first chat ended up blocked. "1년간" (over 1 year), "1종목" (1 stock) and "3종목" (3 stocks) were rejected as quantities. A period like "1-year volatility" isn't a ratio, an amount or a quantity, so durations went into the allowlist. Stock counts were a place where the tools gave no placeholder, leaving the model no choice but to write a number, so the chat prompt now says to list stocks by name instead of counting them.
Summary
With each fix I tried not to loosen the checks. Made-up citations and directly written numbers are still blocked; only cases where the same thing was written differently, or the check miscounted, were let through.
Looking back, most false positives weren't the model's fault. Either the prompt and the check described different rules (years, sentence counts), or the check didn't know what real output looked like (bracketless footnotes, footnotes after periods). So the work started with counting rejection reasons and reading the flagged sentences.