AI AgentsAI Processing of Incident Alerts, with Output Verified by Code
· PinLog
- #observability
- #alerting
- #hermes
- #ai-agent
When adding AI to incident notifications, I had a question to settle before worrying about good prose: which decisions should AI be allowed to make? Turning an alert into readable text is not the same as notifying an entire channel. Changing a sentence and changing a mention policy carry different weight.
In PinLog, Sentinel, which receives and forwards alerts, uses Hermes, an AI agent tool, to process alert payloads (the raw data an alert carries). But the AI's output is not automatically the final message. Trusted code revalidates mentions, URLs, and formatting.
Problem
Adding AI to incident alerts risked handing generated output not only the explanation, but also who gets notified and which links and format go out. Operational policy would then depend on generated output.
Decision
Hermes processes the alert payload, and trusted code revalidates mentions, URLs, and formatting. Notification and repeat policies by severity and state, and the place of the closing one-line summary, are fixed as policy.
Result
The broad question “do we trust AI?” became the concrete question “which parts of its output must code check?” Correct format still does not guarantee correct content.
Alerts Explain and Summon
The person receiving an operational alert needs to know what happened, whether it is still happening, and what to do now. That is why a payload is turned into something a person can read. A message contains more than an explanation, though. A mention requests someone's attention, and a URL points to the next place to look.
I did not want to treat all of these as free-form text generation. If freedom to rephrase also meant freedom to change who gets notified or which message rules apply, operational policy would depend on generated output. Before getting a plausible message for each alert, I wanted the same policy to apply to the same state.
Hermes Processes, Code Decides
Alerts raised inside the cluster flow like this. The monitoring tool Prometheus raises the alert, and Alertmanager hands it to the Sentinel Receiver on the host. Sentinel then delivers it to Mattermost, the team messenger. Sentinel uses a dedicated Hermes profile for alerts. What matters is not the profile's name, but the separate code validation after processing.
Prometheus
Alert originatesAlertmanager
Alert is forwardedSentinel Receiver
Hermes processing · code revalidationMattermost
Operational message delivery
Within that flow, Sentinel's responsibilities split as follows. Rather than hiding the AI stage, the design makes clear where its output meets the operational rules.
Hermes's responsibility
Process the alert payload, taking part in preparing a message for people to read.
Trusted code's responsibility
Revalidate mentions, URLs, and formatting in the processed output. Generated text is not the authority on the final format.
Compared with sending only fixed text, this adds a stage that rechecks the processed output. Compared with sending generated output as-is, it gives AI less discretion. I chose to accept that restriction: leave room to process the explanation, without handing over the rules of an operational message.
Trusted code is not a device that makes every AI explanation true. It is a boundary around who gets notified and which links and format go out. Natural-sounding prose alone is not a reason to treat that boundary as passed.
Severity and State as Policy
The notification policy distinguishes Critical from Warning, and FIRING from RESOLVED. FIRING means the alert is active; RESOLVED means it has cleared. Repeat interval and mention behavior should be read as separate conditions, not lumped together.
Critical FIRING
One
@channelmention, with a one-hour repeat interval for notifications.Warning FIRING
No mention, with a six-hour repeat interval.
RESOLVED
The policy is to always send the resolution notification.
I set Critical to page with one @channel mention and repeat hourly, and Warning to repeat every six hours with no mention. How mentions behave on repeated messages, and how one incident's duplicates are grouped, are outside this article. The point is that severity decides both how people are called and how often alerts repeat.
A resolution notice matters because an initial alert alone does not tell the reader the incident's current state. The policy keeps someone who saw the warning from having to guess whether it cleared. Here, “always send” means resolution is included in what gets sent. It is not a guarantee of arrival regardless of the network or the receiving system.
Three Facts in the Final Line
The final line of an operational alert is a one-line summary of what happened, the current state, and the required action. It is not decoration to end the message neatly. It gives the reader a fixed place to find what they need at the end.
Occurrence · current state
Read what happened together with the state now. The fact that something occurred should not imply it is still ongoing.
Required action
Identify what to do after reading the explanation. This separates an actionable summary from one that only describes the situation.
Writing only “what happened” leaves out the current state. Writing only “is it healthy now?” loses the event behind the message. Including the required action connects the explanation to the reason for reading the alert. AI processing does not change the purpose of that final line.
Format Control Is Not Fact
Looking back, what mattered was not simply putting AI in the alert path. It was the responsibilities that stayed outside it. Separating payload processing from operational rules turned the broad question “do we trust AI?” into the more concrete question “which parts of its output must code check?”
Still, format correctness and factual correctness are different. Mentions, URLs, and message structure can follow the rules while the explanation or required action does not match the real situation. This boundary should not be mistaken for a guarantee of accurate root-cause diagnosis.
The principle I am taking forward is simple: let AI process information, but keep responsibility for calling people's attention and enforcing the final message rules in code. Adding AI to notifications was not only about delegating more decisions. It also meant deciding first which decisions not to delegate.