We wired an agent into every production alert: OpenTelemetry into SigNoz, tools over MCP, a model we host ourselves — no SaaS, no vendor API. It rarely found a root cause, and the alerts were to blame, not the model: a rotation paging monthly, restarts that were counter resets, gaps between rules. Now it grades every alert — cutting what stays green, drafting what incidents proved missing, flagging what nobody can explain. You will leave knowing which rules to delete and which you lack.
AIEEV
Site Reliability Engineer
Victor (Changyun Lee) is a DevOps/SRE engineer at AIEEV, a startup based in South Korea. Starting his career as a frontend and backend developer, he transitioned into infrastructure and reliability engineering. He loves keeping up with the latest tech trends and is currently diving into LLM engineering.