AI AgentsSeparating an Operations AI’s Memory from Actual Server State
· PinLog
- #Pico
- #Knowledge Graph
- #Operations
As the infrastructure owner for PinLog, I worked with information that carried different time horizons. The branch that Argo CD, our deployment tool, deploys from is configuration. Whether a Pod (the unit Kubernetes runs) was Ready is a state observed at one particular moment.
Pico, our operational AI, needed both kinds of knowledge. But they could not be trusted in the same way. Reading yesterday's healthy status as today's can lead to a wrong operational decision, even when the record itself is accurate.
Problem
Pico's operational knowledge mixed structures verified in configuration with states observed at particular moments. Reading a record of past health as the current state could lead to a wrong operational decision.
Decision
Observations and implementation contracts were read separately, and source evidence (Raw), knowledge history (events), and query data (current) were given distinct roles. The Pico knowledge tool stayed read-only, while ingestion and mutation happened in a separate repository.
Result
The principle that remained: use memory to find what to inspect, and verify facts needed for current decisions in the actual system. Read-only access still leaves room to misread historical material.
Rules and State Age Differently
When reading operational information, it helps to separate observations from implementation contracts. An observation is a value inspected at a particular moment. A contract is a structure verified in repository configuration and operating documents.
Observation
The monitoring workloads (the services running for monitoring) were Ready.
This describes their state then. Knowing their health now requires another query.
Implementation contract
Argo CD reconciles deployments against main.
This describes the operating structure. Later configuration changes still need checking.
A contract is not timeless either. It is simply a different kind of information from health data that can change from moment to moment. When reading a search result, the first question was not whether it sounded plausible. It was what it described, and as of when.
Source, History, Query Data
Pico's memory lived in Pico Graph. Pico Graph was not a wiki for people to read. It was an internal knowledge store that Pico searched by source and relationship. Its data had distinct roles.
Immutable Raw
Captured source evidenceevents
Append-only canonical knowledge historycurrent
A rebuildable query view derived from history
This diagram summarizes the roles of the data. It does not imply real-time synchronization with the running server.
store/events.jsonl held the canonical knowledge history, while store/current.jsonl was a rebuildable query view. The word "current" should not be confused with the current operating state. Even an up-to-date query view can contain old observations.
Keeping sources separate from query results gives a route back to the evidence behind an explanation. But having a source and still being valid today are different things. Source and time had to be read together.
Separate Read and Write Rights
Pico's knowledge tool, opened up for use in Hermes, could only read. Collecting and changing knowledge were explicit steps done in a separate repository.
So finding an operational fact was not the same operation as changing canonical knowledge. A guess made while composing an answer could not rewrite the operating record directly through the retrieval tool.
Read-only access did not make every answer accurate, though. Misreading historical material or dropping its date was still possible. Separating write permissions was no substitute for accuracy.
"Was Healthy" vs. "Is Healthy"
Take the question, “Is monitoring healthy now?” A past Ready observation is evidence about the moment it was recorded. Using it as-is to answer a question about now skips a necessary check.
Identify the time horizon
Distinguish a question about historical structure from one about current state.
Use memory to identify what to inspect
Use recorded structure and evidence to narrow down the systems and fields to query.
Verify current decisions at the source
Inspect the actual host, Kubernetes, GitHub, or other source system. If inspection is unavailable, disclose that only historical evidence is available.
These three steps are the rule I set for using memory. Not every question automatically triggered a live check. Operational knowledge was better used to guide current checks than to replace them.
Memory Is Not the Server
Looking back at the record, what mattered was the role of each piece of information, not how much was stored. Raw preserved evidence, history recorded how knowledge changed, and query data supported retrieval. Answers about the current state still had to come from the actual system.
Giving an AI memory helps it recover past context. But that memory does not replace the server as it is now. Use memory to find what to inspect, then verify the facts the current decision needs. That is the principle I want to keep from Pico's operational knowledge.