EngineeringCutting LLM Costs with Caching and a Morning Batch
· FINCH
- #llm
- #cache
- #cost
- #performance
FINCH's AI features are stock analysis, a daily briefing, portfolio diagnosis and chat. All of them call an LLM, and the LLM credits we could use were fixed. In mid-September the infra part counted what was left: 52,745 credits. At the rate we were spending, that was 24 analyses, or 1.6 days. The demo was on September 28.
This post covers how we cut calls and tokens to make it to the demo. Along the way, the cache turned out to behave differently from what we intended, twice.
First, where the tokens went
The same count showed that 68% of tokens went to regeneration. FINCH checks LLM output automatically and regenerates with the reasons attached when a check fails. The regeneration limit and the cache lifetime were hardcoded, so tightening them meant a deploy. I moved both into settings and tightened them right away.
- Regeneration limit: 3 down to 2 (one first attempt plus one retry)
- Cache for shared stock analysis sections: 6 hours up to 24
There was a request to go down to 1, but I kept 2. Those stats were from before a false-positive fix in the checks, so the rejection rate had likely dropped already, and with no retry a rejected section goes out empty. Since it's a setting now, it can go to 1 without a deploy once the new rejection rate is measured.
Cutting reasoning tokens
Next were reasoning tokens. A reasoning model spends tokens thinking before it answers, and the effort setting controls how much. With the same prompt, high took 28 seconds, medium 16 and low 4, and the output was the same. A 90-token answer came with 4,480 reasoning tokens.
28s
effort high
16s
effort medium
4s
effort low
The default went to low on every path, and the output limit went from 16,000 to 4,000 tokens. Personalized sections now reuse the earlier response for the same user, stock and ledger date, and the briefing batch only covers users active in the last 7 days instead of 30.
Two ways the cache didn't work as intended
Looking at production logs the same day, I found two problems in the cache itself.
One failed section wiped out the rest. Stock analysis builds seven sections in parallel. If one timed out, the whole request ended in a 504 and the finished sections were thrown away too. And if one section was empty because it failed a check, the cache lookup discarded the whole row, so that stock regenerated all seven sections on every request. Now a failed section comes back empty, and the cache takes the latest successful result for each section.
The cache never expired. Cache hits were also written to the response log, and the next lookup picked those entries up as "responses from the last 24 hours", so the expiry kept moving forward. In production, a Samsung Electronics analysis generated on September 14 with 0 sources had been copied for two days, and 94 documents ingested in the meantime went unused. Now only entries that were actually generated (cached=false) count as a cache source.
Morning batch instead of user requests
Generating on request makes the user wait, and sometimes builds the same thing more than once. So the heavy generation moved to a batch that runs every morning at 9:20.
Portfolio diagnosis
Diagnosis was built on the first request, so the user waited, and the cache key (a fingerprint) included the current price, so it regenerated whenever prices moved during the day. The 7-day cache hit rate in production was 26%. Taking the price out of the fingerprint made it once a day, and the morning batch now builds diagnosis along with the briefing. Requests return the stored diagnosis; only a first-time user with nothing stored gets one built on request.
Retrying failed sections
When a batch-built stock analysis had a section that failed a check, that section was empty, and the first person to open the stock waited for the retry. Two sections retrying twice each took over ten seconds. The request path no longer rebuilds them; the next morning's batch retries instead.
Daily batch token cap
One pass over 30 stocks takes about 600k tokens. On September 16 a prompt change emptied the cache, the 2M daily cap ran out, and the batch stopped after 2 of 30 stocks. At low effort, 5M tokens cost about $1, so the cap went up to 5M.
When the cache hid a measurement
In late September, while looking into performance-attribution summaries being rejected by the checks, I ran each combination of effort and regeneration limit 3 times. Runs 2 and 3 hit the cache and never called the LLM. A result I read as 0/3 was really 0/1. Responses with a rejected summary were also cached for 10 minutes, so any attempt to reproduce the failure was immediately served from cache.
At the time I got around it by overriding an internal function at runtime on the production pod, which isn't a proper way to check anything. So the function got an option to skip the cache. It's not exposed over HTTP, though: if users could force regeneration, it would be a way around the token budget.
The bottleneck was the number of calls, not the model
On the morning of the demo there was a request to make chat more responsive, so I measured it in production. "How's my portfolio?" took 13.8 seconds with 2 LLM calls. "Why did Hyundai Motor drop?" took 27.3 seconds with 3 calls and ended up blocked. The time came from the number of calls and the output length, not the model size.
So instead of a bigger model, I switched to one with shorter time per call (gpt-6-luna) and turned off reasoning on the turns that pick tools. This model returns a 400 error when tools and reasoning are used together, so just changing the model name would have broken chat immediately. The final call that writes the answer keeps low effort.
13.8 → 8.2s
How's my portfolio?
27.3 → 8.4s
Why did Hyundai Motor drop? (blocked → answered)
Summary
The cost cuts themselves were simple: fewer regenerations, less reasoning, no rebuilding the same thing, and heavy generation done when nobody is waiting. What took longer was checking that the cache actually behaved that way. Partial failures, a cache that never expired and a cache that hid measurements only showed up once I counted the numbers directly.