<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0"><channel><title><![CDATA[Notes by Doctor Droid]]></title><description><![CDATA[Doctor Droid team shares product guides, demos and best practices in observability.]]></description><link>https://notes.drdroid.io</link><image><url>https://cdn.hashnode.com/res/hashnode/image/upload/v1723284089567/30d18ee5-02b3-42cd-a3d6-b53a35289f14.png</url><title>Notes by Doctor Droid</title><link>https://notes.drdroid.io</link></image><generator>RSS for Node</generator><lastBuildDate>Fri, 02 Oct 2026 17:53:30 GMT</lastBuildDate><atom:link href="https://notes.drdroid.io/rss.xml" rel="self" type="application/rss+xml"/><language><![CDATA[en]]></language><ttl>60</ttl><item><title><![CDATA[Open Index: A Structured Context Layer for AI Agents]]></title><description><![CDATA[We’ve been working on a problem that kept showing up while building AI agents: managing domain context.
MCP gives an agent access to tools. Skills help define how it should behave. Memory can store in]]></description><link>https://notes.drdroid.io/open-index-a-structured-context-layer-for-ai-agents</link><guid isPermaLink="true">https://notes.drdroid.io/open-index-a-structured-context-layer-for-ai-agents</guid><category><![CDATA[Open Source]]></category><category><![CDATA[AI]]></category><category><![CDATA[memory]]></category><dc:creator><![CDATA[Pratik Mahalle]]></dc:creator><pubDate>Sun, 09 Aug 2026 06:20:02 GMT</pubDate><enclosure url="https://cdn.hashnode.com/uploads/covers/6a0febdbd8e265f60d9348f4/8c0e7c2a-5547-4c40-b5a4-2bdf78e68ecd.jpg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>We’ve been working on a problem that kept showing up while building AI agents: managing domain context.</p>
<p>MCP gives an agent access to tools. Skills help define how it should behave. Memory can store information from previous interactions.</p>
<p>But a domain-specific agent also needs to know how the things in its world fit together.</p>
<p>A support agent needs to understand customers, accounts, tickets, products and known issues. An SRE agent deals with services, databases, deployments, incidents, and runbooks. A sales agent has companies, people, opportunities, meetings, and decisions.</p>
<p>We built <a href="https://github.com/DrDroidLab/open-index"><strong>Open Index</strong></a> to manage this layer.</p>
<p>It’s an <strong>open-source</strong> framework for representing domain knowledge as structured entities and relationships that an agent can navigate.</p>
<h2>How did we actually land here?</h2>
<p>Our earlier knowledge setup used embeddings and Markdown files.</p>
<p>It was a good place to start. As the amount of knowledge grew, though, we started seeing problems that retrieval alone wasn't solving.</p>
<p>The first was <strong>non-determinism</strong>.</p>
<p>We could replay the same or a very similar scenario and see the agent pick completely different documents. That changed the investigation path and sometimes the final result.</p>
<p>Then there was <strong>navigation</strong>.</p>
<p>A retrieved document might tell the agent something about a service, but the agent still had to figure out what to inspect next.</p>
<p>Which database does this service use? What else depends on that database? Which deployment touched the service recently? Have we seen a similar incident before?</p>
<p>That information existed, but the connections between those pieces weren't always explicit.</p>
<p>We also started seeing <strong>context poisoning</strong>. As more documents and information were added, the agent had more opportunities to pick up something irrelevant or outdated and follow the wrong path.</p>
<p>Updating the knowledge base became difficult too after a certain amount of knowledge.</p>
<p>When the agent learns something new, you need to know where that information belongs, whether it already exists, whether an older version should be replaced, and whether it conflicts with something else.</p>
<p>A folder of Markdown files gets difficult to manage once both humans and agents are continuously reading and writing to it.</p>
<h2>From Markdown to structured context</h2>
<p>We started adding indexes and structured data around our existing context.</p>
<p>That fixed some problems, but created a new set of things we had to handle:</p>
<ul>
<li><p>updates and de-duplication</p>
</li>
<li><p>stale or conflicting information</p>
</li>
<li><p>relationships between entities</p>
</li>
<li><p>version history</p>
</li>
<li><p>tracking which context the agent used</p>
</li>
</ul>
<p>Eventually, this became its own context layer.</p>
<p>We turned that work into <strong>Open Index</strong>.</p>
<h2>How Open Index works</h2>
<p>Open Index lets you define the objects that exist in your domain and the relationships between them.</p>
<p>For example, an SRE setup could have:</p>
<p><strong>Service → Database → Deployment → Incident → Runbook</strong></p>
<p>A support setup could have:</p>
<p><strong>Customer → Account → Ticket → Product → Known Issue</strong></p>
<p>You define the models and schemas. You decide how information gets added or updated. You define how entities are connected.</p>
<p>The agent can then use those connections while working through a task.</p>
<p>Say an investigation starts with a service.</p>
<p>Instead of searching the entire knowledge base every time it needs more information, the agent can follow known relationships from that service to its dependencies, deployments, previous incidents or runbooks.</p>
<p>The context itself gives the agent a map of where it can go next.</p>
<h2>What Open Index handles</h2>
<p>There are three parts we care about.</p>
<ul>
<li><p><strong>Structured memory</strong>: Domain knowledge is stored as structured entities rather than only documents or chunks. The structure is defined by the domain. Open Index doesn't require every agent to use the same schema.</p>
</li>
<li><p><strong>Navigation</strong>: Relationships between entities give the agent paths through the context. An agent investigating a service can move through the objects connected to that service instead of repeatedly trying to rediscover those connections through retrieval.</p>
</li>
<li><p><strong>Relationships</strong>: The connections between entities can represent things such as dependencies, ownership, decisions and correlations. That gives the agent information about how two things are related, not only whether their text happens to be similar.</p>
</li>
</ul>
<h2>Getting started with Open Index</h2>
<p><strong>Open Index</strong> is <strong>domain agnostic</strong>. You can define the entities and relationships based on whatever your agent works with: marketing, sales, legal, support, security, SRE, or something else.</p>
<p>The basic flow is:</p>
<blockquote>
<p><strong>Define your entities and schemas → populate and update context → connect entities → let the agent navigate it.</strong></p>
</blockquote>
<p>You can run Open Index with a single command and see how the setup works before building it into an agent.</p>
<p>The tool is <strong>open source</strong>:</p>
<p><strong>github.com/DrDroidLab/open-index</strong></p>
<p>If you're working on one, we'd like to hear what your context setup looks like and what breaks as it grows.</p>
]]></content:encoded></item><item><title><![CDATA[Kubecon India 2026 Recap with DrDroid]]></title><description><![CDATA[I walked into Jio World Convention Centre on June 17 expecting the usual mix of swag, keynotes, and hallway small talk. Two days later, what stuck with me wasn't any single talk. It was a shift in the]]></description><link>https://notes.drdroid.io/kubecon-india-2026-recap-with-drdroid</link><guid isPermaLink="true">https://notes.drdroid.io/kubecon-india-2026-recap-with-drdroid</guid><dc:creator><![CDATA[Pratik Mahalle]]></dc:creator><pubDate>Wed, 05 Aug 2026 05:41:10 GMT</pubDate><enclosure url="https://cdn.hashnode.com/uploads/covers/6a0febdbd8e265f60d9348f4/c7c8f912-8904-4bde-900d-38c40117ce3c.jpg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>I walked into Jio World Convention Centre on June 17 expecting the usual mix of swag, keynotes, and hallway small talk. Two days later, what stuck with me wasn't any single talk. It was a shift in the conversation itself: cloud native is moving past "how do we observe everything" and into "how do we get something to act on what we observe."</p>
<p>I was there for KubeAutoDay on June 17 and KubeCon + CloudNativeCon India on June 18 and 19, representing DrDroid at the booth. Here's what stood out, where I think things are headed, and what's worth looking into next.</p>
<img src="https://cdn.hashnode.com/uploads/covers/6a0febdbd8e265f60d9348f4/d68690fb-8834-486a-aa18-77d33c9f8137.jpg" alt="" style="display:block;margin:0 auto" />

<h2>KubeAutoDay: the warm-up nobody expected</h2>
<p>KubeAutoDay ran for the first time in India this year, the evening before KubeCon proper started. It felt less like a standalone conference and more like the pre-party: smaller room, more intimate conversations, people comparing notes before the main event got loud. We sponsored it the same way we sponsored KubeCon, but the tone was different. Less pitch, more "okay, what's everyone actually building?"</p>
<img src="https://cdn.hashnode.com/uploads/covers/6a0febdbd8e265f60d9348f4/14a62698-1685-42ca-b8f0-188bc549da0d.jpg" alt="" style="display:block;margin:0 auto" />

<h2>What we saw: two days of booth conversations</h2>
<p>Most of my two days at KubeCon were spent doing two things: getting the booth ready before doors opened, and then standing at it, explaining what DrDroid does to whoever walked up.</p>
<p>That second part is where the real signal was. I wasn't sitting in sessions hearing what speakers think comes next. I was hearing it straight from engineers and platform leads who stopped by, asked what we do, and then told me what they're actually trying to solve right now. A lot of those conversations circled back to the same place: people aren't asking "how do we monitor more," they're asking "how do we get something to act on what we're already monitoring." Automation and AI-driven response came up constantly, not as a buzzword, but as the next thing teams are actively trying to figure out how to adopt.</p>
<h2>Where we're going</h2>
<p>That's exactly the bet DrDroid is making. We're not trying to give teams another dashboard. We're building toward incident response that doesn't wait for a human to start digging: auto-investigation, service discovery, an MCP server that lets agents reason across incident data. The conversations at the booth this year confirmed it. Almost everyone who stopped by wasn't asking if this direction makes sense. They were asking how soon they could get it.</p>
<h2>Good topics to talk about next</h2>
<p>A few things came up enough times at the booth that they're worth writing about properly:</p>
<ul>
<li><p><strong>What teams actually mean when they say "AI agents for incident response."</strong> The phrase gets thrown around a lot. The expectations behind it vary more than people admit.</p>
</li>
<li><p><strong>Where automation should stop and a person should take over.</strong> Every conversation eventually got here, and nobody has a clean answer yet.</p>
</li>
<li><p><strong>What adopting OpenTelemetry looks like for a team that's mid-migration, not greenfield.</strong> Most people we talked to are somewhere in the middle, not starting from scratch.</p>
</li>
<li><p><strong>What changes for a platform team once incident response is built in from day one</strong>, instead of added after the first bad outage.</p>
</li>
</ul>
<h2>Find us there next time</h2>
<p>If you stopped by the booth this year, thank you for the conversation. If you didn't, we'll be back at the next one. Come find us, skip the small talk, and tell us about your last incident. That's the conversation we actually want to have.</p>
]]></content:encoded></item><item><title><![CDATA[What Is AI SRE?]]></title><description><![CDATA[Every vendor has a different answer. Here's ours, and the test we use to tell the real thing from a chatbot with a status page.
Type "AI SRE" into a search bar and you'll get four different answers, n]]></description><link>https://notes.drdroid.io/what-is-ai-sre</link><guid isPermaLink="true">https://notes.drdroid.io/what-is-ai-sre</guid><category><![CDATA[AI]]></category><category><![CDATA[SRE]]></category><category><![CDATA[Devops]]></category><category><![CDATA[Platform Engineering ]]></category><dc:creator><![CDATA[Pratik Mahalle]]></dc:creator><pubDate>Tue, 04 Aug 2026 06:47:30 GMT</pubDate><enclosure url="https://cdn.hashnode.com/uploads/covers/6a0febdbd8e265f60d9348f4/f79b5690-12ef-4bfe-a64a-5d12fa44b4b1.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><em>Every vendor has a different answer. Here's ours, and the test we use to tell the real thing from a chatbot with a status page.</em></p>
<p>Type "AI SRE" into a search bar and you'll get four different answers, none of which agree with each other. One vendor means "an AI chatbot bolted onto your alerts." Another means "we replaced your on-call rotation with an agent." A third just means "we added an LLM summary field to our dashboard." Same two words, three unrelated products.</p>
<p>That's not a branding problem. It's a definition problem, and it's worth fixing now, before "AI SRE" hardens into shorthand for whoever raised the biggest round instead of whoever built the thing that actually diagnoses an incident.</p>
<h2>A Working Definition: Software That Performs the Work</h2>
<p>AI SRE is software that performs the diagnostic and remediation work of a site reliability engineer during an incident: correlating signals across your stack, forming a hypothesis about root cause, and either acting on that hypothesis or handing a human a ready-to-verify answer.</p>
<p>The key word is <em>performs</em>, not <em>assists</em>. A tool that summarizes an alert for you is assisting. A tool that pulls logs, traces, and deploy history, rules out three plausible causes, and tells you "the checkout service started erroring at 14:32, right after the payment-gateway config push, here's the diff" is performing SRE work. Those are different products even when they share a chat interface.</p>
<p>That distinction is the whole ballgame, and most of the confusion in the space comes from vendors blurring it.</p>
<h2>Why the Definition Keeps Shifting</h2>
<p>Three things are colliding at once, and none of them are fully sorted out.</p>
<p><strong>The label got attached before the behavior did.</strong> "AI SRE" started showing up in pitch decks around the same time GPT-4-class models made it plausible to reason over structured telemetry data. Plenty of companies had the label ready before they had the reasoning loop working. So the term picked up baggage from products that never closed the diagnostic loop, they just added a natural-language layer over an existing dashboard.</p>
<p><strong>AIOps never fully died, and its scope overlaps.</strong> AIOps promised anomaly detection and noise reduction: fewer alerts, smarter routing, some correlation. AI SRE promises something further downstream, actual root cause reasoning and remediation. But because both terms touch "AI plus operations," people use them interchangeably, and vendors let them, because AIOps has a decade of enterprise trust behind it that a two-year-old category doesn't.</p>
<p><strong>The agent part is genuinely unsettled.</strong> Does AI SRE mean a fully autonomous agent that restarts pods and rolls back deploys without asking? Or does it mean a highly capable co-pilot that never takes write actions, only read actions plus a recommendation? Both exist in the market right now under the same name, and the risk tolerance for each is wildly different. A bank's platform team and a five-person startup have opposite answers to "should this thing touch prod," and the category hasn't split into those two lanes yet.</p>
<h2>Three things AI SRE gets confused with</h2>
<p>It helps to draw the boundary by exclusion.</p>
<p>It's not a chatbot wrapper on top of your existing monitoring stack. If the product's core value is "ask questions about your dashboards in English," that's a UX improvement on observability, not SRE work. SRE work means correlating across tools your dashboard doesn't already connect: the deploy pipeline, the incident channel, the runbook, the on-call schedule, the actual code diff.</p>
<p>It's not anomaly detection with a new name. Flagging that a metric moved 3 standard deviations from baseline is useful and has existed for years. It's not diagnosis. Diagnosis is explaining <em>why</em> it moved and what to do about it.</p>
<p>It's not a summarizer. Taking a wall of log lines and shrinking them into three sentences saves time, but it doesn't reduce mean time to resolution on its own unless the summary points at a cause a human wouldn't have found as fast.</p>
<h2>A working test you can apply to any claim</h2>
<p>Here's a rough test for whether something qualifies: can it answer "why did this break" without a human first telling it where to look? If a human has to point the tool at the right service, the right time window, and the right dependency before it produces anything useful, it's a search tool with a nice UI. If it can start from an alert and narrow the blast radius on its own, using context it gathered rather than context it was handed, that's closer to the real thing.</p>
<p>None of this means the category is fake. It means it's early, the same way "cloud native" was messy in 2015 before CNCF gave it edges. The underlying need, cutting the time between "something's wrong" and "here's why and here's the fix," is real and it's getting worse as systems get more distributed and more people ship code without ever learning the infrastructure underneath it.</p>
<h2>Where DrDroid draws the line</h2>
<p>We built <a href="https://drdroid.io/">DrDroid</a> around the "performs, not assists" side of that test, so it's worth saying plainly where we land on the questions above rather than leaving it implied.</p>
<p>On read versus write: DrDroid's default posture is diagnostic. It correlates across your alerting, logs, traces, deploys, and infra config on its own, forms a root cause hypothesis, and hands it to your on-call engineer to verify, rather than acting unprompted in prod. Teams can wire up remediation actions on top of that once they trust the diagnosis, but trust has to come first. An agent that's confidently wrong and empowered to restart services is worse than no agent at all.</p>
<p>On the AIOps overlap: AIOps tools tell you <em>something changed</em>. DrDroid's job starts one step later, explaining <em>what changed and why it matters</em>, using the same postmortem logic an experienced SRE would apply: what deployed recently, what's upstream and downstream of the failing service, what looks different from last week's baseline.</p>
<p>That's also why "does it need a human to point it at the problem" is the test we hold ourselves to internally, not just a line for a blog post. It's the difference we're trying to build toward, and it's the reason we'd rather argue for a tight definition of this category than a loose one that includes us by default.</p>
<h2>Where the Category Is Heading</h2>
<p>Expect the term to fracture into sublabels over the next year or two, the same way "observability" split into logs, metrics, and traces once everyone agreed on what the umbrella term meant. My guess: something like "diagnostic AI SRE" (reasoning, read-only, human approves the fix) and "autonomous AI SRE" (write access, acts within guardrails) become the actual distinction people care about, and "AI SRE" alone stops meaning much on its own, the way "cloud" alone doesn't tell you if someone means SaaS, IaaS, or a Raspberry Pi in their closet.</p>
<p>Until then, the honest move for anyone building or buying in this space is to ask the specific question instead of trusting the label: does it read or does it write, does it need a human to point it at the problem or does it find the problem itself, and what happens the first time it's wrong.</p>
]]></content:encoded></item><item><title><![CDATA[The Roadmap And Path To Autonomous AI in Software Monitoring]]></title><description><![CDATA[The difference between the future and the status quo of production monitoring operations is vast.
The Future:
At one end, a human reads alerts, opens issues, and fixes manually. At the other end is th]]></description><link>https://notes.drdroid.io/the-roadmap-and-path-to-autonomous-ai-in-software-monitoring</link><guid isPermaLink="true">https://notes.drdroid.io/the-roadmap-and-path-to-autonomous-ai-in-software-monitoring</guid><category><![CDATA[AI]]></category><category><![CDATA[SRE]]></category><category><![CDATA[#AIOps]]></category><category><![CDATA[automation]]></category><dc:creator><![CDATA[Pratik Mahalle]]></dc:creator><pubDate>Mon, 03 Aug 2026 07:12:58 GMT</pubDate><enclosure url="https://cdn.hashnode.com/uploads/covers/6a0febdbd8e265f60d9348f4/e22d311e-40e9-4f8a-ac16-17956d0c19c1.jpg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>The difference between the future and the status quo of production monitoring operations is vast.</p>
<p>The Future:</p>
<p>At one end, a human reads alerts, opens issues, and fixes manually. At the other end is the future, where humans are rarely involved in Day 2 operations; a system finds the problem, works out the cause, and fixes itself, while a human is only tagged in when needed.</p>
<p>The useful question there isn't "are we automated?" It's like: what specifically is stopping us from moving up one level? And is there any merit in moving another layer up?</p>
<p>At DrDroid, we've been iterating over the different layers of automation &amp; autonomy in production monitoring systems.</p>
<p>In this blog, we have discussed a model to define different levels of autonomy in software systems and what each of these levels means.</p>
<img src="https://cdn.hashnode.com/uploads/covers/6a0febdbd8e265f60d9348f4/0f527814-b7da-4a90-be80-beeb421da590.jpg" alt="" style="display:block;margin:0 auto" />

<h2>Level 0: Observability</h2>
<p>Observability is the first layer of moving towards autonomous infrastructure. Here's what it means for a software system to be classified as a Level 0 system:</p>
<ul>
<li><p><strong>Data exists:</strong> Your stack is instrumented with fair coverage of logs/metrics/traces/events so that unknowns can be identified and resolved (by humans).</p>
</li>
<li><p><strong>Signals exist:</strong> Exhaustive coverage of alerts across infrastructure, application, and CUJs (Critical User Journeys) that help your team identify/detect early on.</p>
</li>
<li><p><strong>Human driven:</strong> A human does everything. An alert tells you something is wrong; you decide what that means, where to look, and what to do about it.</p>
</li>
</ul>
<h3>Level 0: Outcome</h3>
<p>An engineering team has enough data at most points to be able to detect, diagnose, and fix an issue with the data points. And that is where the first level ownership comes in.</p>
<p><strong>Ownership:</strong> Engineers &amp; Teams own everything.</p>
<h2>Level 1: Scripted automation</h2>
<p>This is where most of the Pre-AI era automations reside. For example,</p>
<ul>
<li><p>Have a known issue where pods need to be restarted.</p>
</li>
<li><p>Have a user related alert - auto fetch 5 more data points and append them to the alert.</p>
</li>
</ul>
<p>Folks automated the runbook. Beyond the runbooks, engineers handle everything the scripts don't cover, which is still most of the work to do.</p>
<h3>L1 outcome:</h3>
<p>The biggest value add often of the first layer is</p>
<ul>
<li><p>Fallbacks with quick remediations come into picture.</p>
</li>
<li><p>Automations by engineers who would otherwise manually do the exact same steps.</p>
</li>
</ul>
<p>Scripted automation is reliable and brittle at the same time. It does exactly what you told it and nothing else. The moment an incident doesn't match a known pattern (which is quite often!), you end up in the same loop of humans doing things. Also, they can get outdated as systems evolve.</p>
<p><strong>Tools:</strong> Automation &amp; workflow frameworks like <a href="https://github.com/DrDroidLab/playbooks">Playbooks</a> and StackStorm, or just plain vanilla bash/python scripts are the common strategies used by teams here.</p>
<h2>Level 2: Agentic investigation</h2>
<p>This is the first level where the machine comes into play and does real work instead of being an NPC. For example, an alert fires, and an agent pulls the relevant logs, correlates the recent deploy, checks the downstream service, and hands you a first read on what's going on.</p>
<p>But you still have to decide and act. You don't start things from scratch - the agent compresses the first twenty minutes of every incident into something you can read in a few minutes, instead of manual assessment.</p>
<h3>L2 Outcome:</h3>
<p>When an L2 layer is successfully implemented, here's what teams benefit from it:</p>
<ul>
<li><p>Humans receive an alert along with detailed investigation - Agent identifies what triggered this issue, accurately and consistently.</p>
</li>
<li><p>While humans might have skipped analysis or serialised analysis when there's a surge of alerts, here the agent can run parallel investigations and help get up to speed faster.</p>
</li>
</ul>
<p>Note: A successful Level 2 autonomy layer in your software infrastructure means that the system can auto-identify the triggers <em><strong>accurately</strong></em>. "Accurately" is important because a system used in production infrastructure needs high reliability for engineers to trust in high stakes environment.</p>
<h3>Pre-requisites:</h3>
<p>While this investigation agent sounds simple early on, there are nuances that make it tricky to get to the end line:</p>
<ul>
<li><p><strong>Agent with access to tools:</strong> As you retrospect over more and more production incidents, it's clear that only the logs and metrics don't tell the full picture. Teams often end up checking deployment configs (in CI/CD tools), code review (in GitHub repository), customer level data (in internal tools/databases), security/configs in cloud and more. For an agent to be able to pinpoint triggers better, it needs access similar to the access that a senior engineer might have.</p>
</li>
<li><p><strong>Context:</strong> Tribal knowledge, company infrastructure, and known patterns are a big giveaway when it comes to production troubleshooting. You need to setup a basic layer of Context for the agent to be able to operate well - it could begin, initially, with some architecture docs and some skills, but eventually as you try to iterate and improve the agent's accuracy so it becomes near accurate in every instance, it needs to be a continuously updating context map/knowledge graph about your infrastructure.</p>
</li>
</ul>
<img src="https://cdn.hashnode.com/uploads/covers/63200bf16c86d75accc7fd61/46217875-fb3c-4b9d-8e51-0c105bb06794.png" alt="A representative of what a mental model of context could look like" style="display:block;margin:0 auto" />

<p><strong>Ownership:</strong></p>
<ul>
<li><p>At this stage, tickets &amp; alerts are still assigned to humans. Dedicated engineers (on-call) are still looking at noisy alerts in Slack. The first responder on a Pager is still an engineer.</p>
</li>
<li><p>The engineer does not have the full burden of doing the heavy lifting, but validating the AI's insight &amp; recommendations.</p>
</li>
<li><p>The builder of the AI Agents + service owner has the implicit responsibility to ensure that the Agent has the right context + right harness to avoid hallucinations, poor correlations.</p>
</li>
</ul>
<h2>Level 3: Cognitive decisioning</h2>
<p>At Level 3, we are turning the tables with regards to ownership and workflows. Humans are not dedicatedly looking at Day-2 ops every now &amp; then.</p>
<p>The agent doesn't just investigate an issue after it's identified as an alert or incident. It triages every signal, identifies if it's actually an issue, clusters it with other existing/ongoing issues and routes it to the right owner (with the root cause proposed).</p>
<h3>Level 3 Outcome:</h3>
<ul>
<li><p>The ownership of sifting through alerts &amp; signals --&gt; detecting issues and identifying the owners of the actual trigger of the issue are not done by on-call engineer, but by the Agent.</p>
</li>
<li><p>Agents consume all alerts &amp; signals, run all investigations and cluster them into actionable buckets. Humans review this on demand.</p>
</li>
<li><p>The agent also has access to known fixes and can apply them without approval. It can also apply new fixes but needs peer (human) review, as is the norm with any process today.</p>
</li>
</ul>
<h3>Pre-Reqs:</h3>
<ul>
<li><p><strong>Root cause quality:</strong> If your agent isn't able to investigate well, it will have exponentially higher risk here because the failure would compound across multiple queries, and poor data can lead to partial assessments.</p>
</li>
<li><p><strong>Context Map:</strong> When an agent needs to be able to discover and navigate a given architecture, skills alone wouldn't help.</p>
</li>
<li><p><strong>Self-learning:</strong> The architecture upgrades every day. The agent's context map needs to be able to continuously update itself - learning from everyday changes, patterns and more.</p>
</li>
</ul>
<p>The effect here is that it has what you can call a pseudo world model - where all the known components are known.</p>
<p>What makes this possible is the context graph, which I'd call a <strong>pseudo world model</strong>: a representation of services, dependencies, ownership, and normal behavior that lets the agent reason about consequences instead of pattern-matching on symptoms. It knows that restarting this pod drains that queue and that the queue feeds a job someone cares about.</p>
<p>The word "pseudo" is doing a lot of work here. The model is good enough to propose the right action most of the time, which is why the human role collapses to a single thing: approve. The agent does the thinking; you say yes or no. And nothing gets resolved on the agent's say-so alone, because the model is approximate, and approval is where a person catches the cases it gets wrong.</p>
<p>This is where on-call changes shape. The reason is specific, and it's the part most people get backwards.</p>
<blockquote>
<p>And after level 3, you even don't need human on on-call</p>
</blockquote>
<p>People go on call because someone has to be accountable for acknowledging an incident and seeing it through to resolution. That's the actual job.</p>
<p>Below level 3, that ownership can't be handed off. The machine can investigate, but it can't be trusted to act, so a human has to be reachable to make the call and give the fix.</p>
<p>Here, the shape inverts. The agent triages, routes, and proposes around the clock. A person isn't tagged for every incident anymore. They get pulled in for two cases: the agent got stuck, or it found something critical enough that it wants a human in the loop before making a decision. On-call stops being "wake up for every alert" and becomes "be reachable when the system asks for you."</p>
<p>That's the real reason to climb the ladder.</p>
<h2>Level 4: Autonomous resolution</h2>
<p>This is the highest level in AIOps. The system detects, decides, fixes, and closes the loop on its own; we can also say it is self-healing. Humans provide oversight, not sign-off, on each action.</p>
<p>The gap between level 3 and level 4 is just one word, and that changes the whole game: the model goes from pseudo to true.</p>
<p>A pseudo world model is a good map. A true world model is one which is accurate enough that the system can act on its own predictions without a person checking each one. The difference is trust, and trust here is earned through accuracy on the consequences that actually matter. At level 3, you approve because the model might be wrong. At level 4, you don't, because it's been right often enough, on the decisions that count, that supervision can replace approval.</p>
<p>That's a research problem, not a ship date. Building a model of a live distributed system that's accurate enough to trust with self-healing is hard for the same reason modeling is hard everywhere in AI; the map drifts from the territory the moment you finish drawing it, and a distributed system rewrites its own territory every deploy.</p>
<h2>How We're Building This Future</h2>
<p>At <a href="https://drdroid.io/">DrDroid</a>, we don't think the journey to autonomous operations happens overnight.</p>
<p>The biggest gap today isn't better LLMs; it's reliable context. Production incidents require reasoning across monitoring systems, deployments, cloud infrastructure, code repositories, ownership information, historical incidents, and internal documentation. Without that context, AI agents remain assistants instead of operators.</p>
<p>That's why our focus has been on building AI SRE agents that combine deep infrastructure context with access to operational tooling. Today, those agents help engineering teams investigate incidents automatically, correlate signals across systems, identify likely root causes, and reduce the manual effort required during on-call.</p>
<p>We're also investing heavily in the next stage: creating a continuously evolving infrastructure context graph that enables higher levels of autonomy, including intelligent alert grouping, issue clustering, ownership discovery, and eventually AI-assisted operational decision making.</p>
<p>We don't believe autonomous operations arrive through one breakthrough model. They arrive by solving each layer of the maturity model with enough reliability that engineers naturally trust the next one.</p>
<h2>So, where are the teams actually?</h2>
<p>Most teams are still at ground level. A growing number are reaching for agentic investigation, and that reach is realistic today. Cognitive decisioning is the current frontier, where the context graph/knowledge graph is the thing separating the teams that talk about it from the teams shipping it, actually.</p>
<p>Autonomous resolution is where organisations wanted to go, but very few products are actually there today.</p>
<blockquote>
<p>The maturity model isn't about measuring how "AI-powered" your operations are. It's about identifying the single capability that unlocks the next stage. Progress isn't made by adding another LLM or another dashboard, it's made by removing the biggest problem in your incident response workflow and that we already know.</p>
</blockquote>
]]></content:encoded></item><item><title><![CDATA[How DrDroid Builds and Maintains the Knowledge Layer That Powers an AI SRE Agent]]></title><description><![CDATA[An AI agent operating on production systems is only as effective as the context it can access at the moment a question is asked. Generic foundation models, however capable, do not know your service na]]></description><link>https://notes.drdroid.io/how-drdroid-builds-and-maintains-the-knowledge-layer-that-powers-an-ai-sre-agent</link><guid isPermaLink="true">https://notes.drdroid.io/how-drdroid-builds-and-maintains-the-knowledge-layer-that-powers-an-ai-sre-agent</guid><category><![CDATA[memory]]></category><category><![CDATA[AI]]></category><category><![CDATA[SRE]]></category><category><![CDATA[Platform Engineering ]]></category><dc:creator><![CDATA[Pratik Mahalle]]></dc:creator><pubDate>Fri, 05 Jun 2026 09:10:45 GMT</pubDate><enclosure url="https://cdn.hashnode.com/uploads/covers/6a0febdbd8e265f60d9348f4/99dd8955-6972-4970-b8bd-cce3fbb9bd6e.jpg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>An AI agent operating on production systems is only as effective as the context it can access at the moment a question is asked. Generic foundation models, however capable, do not know your service names, your dashboards, your deploy history, or your tribal knowledge. Closing that gap is not a model problem. It is a knowledge engineering problem.</p>
<p>This paper describes DrDroid's Context Engine: the continuously updated knowledge layer that captures the state and history of a customer's production environment, and the retrieval system that surfaces the right slice of that knowledge during an agent investigation. We cover what the Context Engine stores, how it is built and kept current, how the agent searches it, and the design choices that make this approach measurably better than RAG, embeddings, or static documentation alone.</p>
<h2>The problem we are solving</h2>
<p>Production engineering is not a single discipline. It is the orchestration of dozens of systems, each with its own data model, vocabulary, and failure mode. A typical mid-sized engineering team operates across:</p>
<ul>
<li><p>A monitoring stack (Grafana, Prometheus, Datadog, Signoz)</p>
</li>
<li><p>A deployment pipeline (ArgoCD, Jenkins, GitHub Actions)</p>
</li>
<li><p>A cloud or hybrid infrastructure provider (AWS, GCP, Azure)</p>
</li>
<li><p>An orchestrator (Kubernetes)</p>
</li>
<li><p>An alerting and on-call layer (PagerDuty, Opsgenie)</p>
</li>
<li><p>A code platform (GitHub, GitLab)</p>
</li>
<li><p>Collaboration and ticketing (Slack, Jira, Confluence)</p>
</li>
<li><p>Databases, APMs, error trackers, analytics tools</p>
</li>
</ul>
<p>When an alert fires, the engineer who resolves it fastest is not the one who reads the alert text most carefully. It is the one who carries an internal map of how these systems connect for this specific company. They know which dashboard is the real one. They know that payments-svc and checkout-payments are the same service under two names. They remember that a similar alert two weeks ago was caused by a Redis eviction, not a code change.</p>
<p>This map is built over months of on-call rotations. It is the difference between a five-minute resolution and a fifty-minute one. And until now, it has been almost impossible to give to an AI agent.</p>
<p>A foundation model knows what Kubernetes is. It does not know what your Kubernetes cluster looks like. RAG over your documentation helps a little, but production reality is not in your documentation. It is scattered across live tools, recent deploys, evolving log patterns, and engineer's heads.</p>
<p>The Context Engine is our solution to this gap. It is the knowledge layer that makes a general-purpose AI agent behave like a senior engineer who has been on your team for two years.</p>
<h2><strong>What the Context Engine is</strong></h2>
<p>A short definition first:</p>
<blockquote>
<p><em><strong>The Context Engine is a structured, continuously updated knowledge layer about a customer's production environment, designed to be queried by an AI agent during operational tasks.</strong></em></p>
</blockquote>
<p>It has two parts. The first is the knowledge itself: a catalog of what exists in the production environment, organized into typed records. The second is the retrieval system that sits on top: a hierarchical search engine, which we refer to as <strong>Dynamic Memory Retrieval (DMR)</strong>, that lets the agent pull the right records at the right moment without dumping everything into a context window.</p>
<p>The analogy we find useful is a library. The knowledge layer is the collection of books on the shelves. DMR is the card catalog and the librarian. You need both. A library without organization is just a pile of paper. A search system with nothing to search is just an empty interface.</p>
<img src="https://cdn.hashnode.com/uploads/covers/6a0febdbd8e265f60d9348f4/578f13d4-33a6-4fbe-a66c-735e0c9c2103.png" alt="" style="display:block;margin:0 auto" />

<p>The Context Engine is a knowledge layer that sits between your production environment and the AI agent. The goal of the layer is simple to state and harder to build: before the agent answers a question about your system, it should already have the same picture in its head that your senior engineer does.</p>
<p>To get there, the engine continuously builds and maintains four kinds of context. I'll walk through each one, what's inside, and how we keep it updated without anyone on your team babysitting it.</p>
<h3><strong>1. The Map of Your Stack</strong></h3>
<p>When you connect DrDroid to your existing tools, the first thing the engine builds is an inventory. Not a CMDB you fill out by hand. An actual working map pulled directly from your APM, ArgoCD, GitHub, and Kubernetes.</p>
<p>This map covers every service with its repos and owning team. Every Grafana and Datadog dashboard, the panels inside them, the queries those panels run. Every database, VM, K8s cluster, and namespace. Every third-party vendor your application talks to.</p>
<p><strong>How we keep it current:</strong> Each source tool gets its own synchronizer. ArgoCD changes get picked up on a webhook. Kubernetes is reconciled on a short interval against the live cluster state. GitHub activity flows in through repo events. We don't trust any single tool to be authoritative on its own, so the engine resolves conflicts using a precedence order per record type. The result is a map that drifts in minutes, not days.</p>
<p><strong>What this changes:</strong> When an alert mentions checkout-api, the agent already knows that means the checkout-api repo on GitHub, owned by the Payments team, running in the prod-us-east cluster, monitored by the checkout-deepdive Grafana dashboard, depending on the orders database and Stripe. It doesn't have to ask. It doesn't have to query five tools to figure it out. It already has it.</p>
<h3><strong>2. How Your Services Actually Connect</strong></h3>
<p>Maps are static. Production isn't. The engine also builds a live correlation layer which has what depends on what, what traffic flows where.</p>
<img src="https://cdn.hashnode.com/uploads/covers/6a0febdbd8e265f60d9348f4/7d671064-0948-4825-a8ce-af94a3ee31b9.png" alt="" style="display:block;margin:0 auto" />

<p>This is where most "AI for ops" tools quietly fall apart, because reconstructing a service graph from a static config is one thing and keeping it accurate while services come and go is another.</p>
<h3><strong>3. Your Team's Tribal Knowledge</strong></h3>
<p>This is the part senior engineers care about most, because this is the part they're tired of being the only source of.</p>
<p>The engine stores runbooks and SOPs, synced from Confluence or Notion or uploaded directly. It stores every past investigation and RCA the agent has run, including what it queried, what it correlated, and what it concluded. It watches your alert patterns over time so it can tell a routine 7 AM spike from a real anomaly. It knows who's on the Payments team and who's on-call this week.</p>
<p><strong>How we keep it current:</strong> Runbooks sync on a fixed schedule and on doc-update events where the source supports them. Investigations are written into the engine as the agent finishes them, with full traceability of which records they touched. Alert pattern fingerprints are recomputed on a rolling window, so a service that used to be quiet and is now noisy doesn't look "normal" forever just because it once was.</p>
<p><strong>What this changes:</strong> A new engineer on-call doesn't need someone senior to tell them "the last three times this fired, it was a config change in service X." The agent says that, because the agent has been watching.</p>
<h3><strong>4. What Changed in the Last Hour</strong></h3>
<p>Most production incidents trace back to something that changed recently. This is the freshest layer of the engine and the one that matters most during a live incident.</p>
<p>It tracks code changes (every PR, commit, and merge across your repos). It tracks infrastructure diffs (new namespace, renamed cluster, database added). It tracks live alerts (what's firing now, what just stopped). And it polls 150+ vendor status pages, so the agent knows if Stripe or AWS is having a bad day before you do.</p>
<p><strong>How we keep it current:</strong> This is the part of the engine that's almost entirely event-driven. PRs and deploys come in on webhooks. Alerts flow in from your alerting tool directly. Vendor status pages are polled aggressively, on the order of every few minutes. The "last hour" delta is recomputed continuously, not cached.</p>
<p><strong>What this changes:</strong> The agent walks into an investigation with a delta of what's different in the last hour. Most of the time, the answer is sitting somewhere in that delta.</p>
<h2><strong>How retrieval actually works</strong></h2>
<p>Building the knowledge is one problem. Pulling the right slice of it at the right time is another. This is the part where most teams trying to build something similar quietly run into a wall.</p>
<p>When we started, the obvious approach was retrieval-augmented generation. Take everything we know about a customer's environment, chunk it, embed it, do similarity search. It's the standard pattern for AI applications, and it works well for general-purpose knowledge bases.</p>
<p>It does not work for production infrastructure.</p>
<p>The reason took us some time to articulate clearly. In production engineering, a single character can mean a completely different thing.</p>
<ul>
<li><p>us-east-1 and us-east-2 are different regions</p>
</li>
<li><p>checkout-prod and checkout-staging are different services</p>
</li>
<li><p>pod-abc123 and pod-abc124 are different pods</p>
</li>
<li><p>v1.4.2 and v1.4.3 might behave totally differently</p>
</li>
</ul>
<p>Embedding models, by design, map semantically similar strings to nearby vectors. They will tell you these pairs are 99% similar. They are not. During an incident, confusing them takes a small problem and turns it into a much bigger one.</p>
<p>The other issue is that production names don't carry semantic content. A service name is an arbitrary identifier. A cluster name is a label. There's nothing for the embedding to "understand." There's just the exact string. Keyword search works much better here, because real production questions almost always anchor on a specific named entity, and that name appears verbatim somewhere in the right record.</p>
<p>So our retrieval layer (we call it Dynamic Memory Retrieval) is keyword-first against the structured graph. We index entities for exact and prefix matching first. We layer embeddings in only for the record types where semantic similarity genuinely helps: runbook prose, ticket descriptions, Slack threads, postmortems. Never as the default for infrastructure records.</p>
<p>A few other design choices worth flagging:</p>
<p><strong>Hierarchical results.</strong> A service ranks above its log lines. A dashboard ranks above its individual panels. A runbook ranks above a casual Slack mention. The agent gets the parent record first and drills into children only when it needs to.</p>
<p><strong>Multi-pass retrieval.</strong> A single search rarely surfaces everything an investigation needs. The agent runs successive queries, each refined by what the previous one returned. A typical end-to-end investigation does between 5 and 50 retrievals against the engine, interleaved with tool calls against the underlying systems.</p>
<p><strong>Short-term and long-term memory, kept separate.</strong> Long-term records describe what's always true: the list of services, the structure of the codebase, the dashboards that exist. Short-term records describe what's currently true: the deploy from 40 minutes ago, the alert firing right now, today's open incident. Mixing them confuses the agent's sense of "now" versus "always." So they're stored and queried as distinct stores.</p>
<h2><strong>How the engine improves over time</strong></h2>
<p>The engine isn't static. Three things keep happening in the background.</p>
<p><strong>It learns from every investigation.</strong> Each RCA the agent produces becomes a record in long-term memory. The next time something looks similar, the agent surfaces it: "This looks like the incident on March 14th where the Redis eviction policy was the cause."</p>
<p><strong>It tracks your environment as it changes.</strong> New services, deleted dashboards, renamed clusters get picked up automatically. You don't re-onboard, ever. The engine that handled your environment on day 1 is materially different from the one handling it on day 90, in a good way.</p>
<p><strong>It stays auditable.</strong> Every recommendation the agent makes is traceable back to which records it pulled and which tool calls it ran. If a senior engineer doesn't trust an answer, they can walk back through the reasoning step by step. This matters more than it sounds like it should, because trust is what determines whether an AI agent actually gets used.</p>
<p>One of our customers described the practical impact like this: "We went from a 90-day onboarding window for new SREs to two weeks." That's the bet, institutional knowledge living in a system instead of in three senior engineers' heads.</p>
<h2><strong>Conclusion</strong></h2>
<p>If you're building or evaluating an AI agent for production work, the single most important question is not which model it's using. It's what context the agent walks into the room with.</p>
<p>A foundation model with MCP servers and no context engine is going to spend the first ten tool calls of every investigation re-learning your environment, and it's going to get half of it wrong because production names look semantically similar but mean entirely different things.</p>
<p>A foundation model on top of a continuously maintained knowledge layer of your services, dashboards, runbooks, dependencies, and recent changes will start every investigation already pointed in roughly the right direction.</p>
<p>That's the whole architectural argument behind the Context Engine. Models will keep getting better. The work of building, maintaining, and retrieving the right context is the part that compounds, and it's the part most teams underestimate.</p>
<p>To understand more about content layer and knowledge layer, check out the youtube videos from @<a href="https://www.youtube.com/@DrDroidDev">drdroiddev</a> channel</p>
<p><em>If you want to see the engine running against your stack,</em> <a href="https://calendly.com/siddarthjain/doctor-droid-discovery-call"><em>schedule a technical evaluation</em></a><em>.</em></p>
]]></content:encoded></item><item><title><![CDATA[How DrDroid’s MCP Server Puts Production Context Inside Claude Code and Any IDE]]></title><description><![CDATA[Your code lives in your IDE. The context you need to debug it lives everywhere else.
That gap is the whole problem. So recently we found the solution for it. DrDroid now has an MCP server, and it puts]]></description><link>https://notes.drdroid.io/how-drdroid-s-mcp-server-puts-production-context-inside-claude-code-and-any-ide</link><guid isPermaLink="true">https://notes.drdroid.io/how-drdroid-s-mcp-server-puts-production-context-inside-claude-code-and-any-ide</guid><category><![CDATA[mcp]]></category><category><![CDATA[claude]]></category><category><![CDATA[context]]></category><dc:creator><![CDATA[Pratik Mahalle]]></dc:creator><pubDate>Tue, 02 Jun 2026 16:19:27 GMT</pubDate><enclosure url="https://cdn.hashnode.com/uploads/covers/6a0febdbd8e265f60d9348f4/2a9d099f-b7b3-4b4b-8aff-0671db77f513.jpg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Your code lives in your IDE. The context you need to debug it lives everywhere else.</p>
<p>That gap is the whole problem. So recently we found the solution for it. DrDroid now has an MCP server, and it puts the production context we've already built for your stack, alerts, traces, past incidents, service relationships, infrastructure metadata, directly inside Claude Code or any MCP-compatible IDE.</p>
<p>If you're already running DrDroid, here's how it works.</p>
<h2>What's the need for it?</h2>
<p>You know the problem. Latency triples, the failing trace is open in your editor, and the fix is probably one line. But the line you need depends on what you can't see from the trace alone. The alert that fired four minutes ago. The deploy that went out at 14:32. The incident from three weeks back with the same signature.</p>
<p>DrDroid already knows all of that. It's in the knowledge graph. The problem was that the graph lived on DrDroid's side and your editor lived on yours, so you'd jump out of the IDE to go get it. The MCP server removes the layer now. The context graph you already rely on now answers from inside the editor where you're writing the fix.</p>
<h2>How it work</h2>
<p>Connect your IDE and you're querying the context graph directly. On the other end:</p>
<ul>
<li><p><strong>Alerts:</strong> current and historical, across every monitoring platform you've connected</p>
</li>
<li><p><strong>Traces:</strong> distributed tracing data from your services</p>
</li>
<li><p><strong>Past incidents:</strong> the full investigation history and root cause analysis, not just that something broke but how it was fixed</p>
</li>
<li><p><strong>Service relationships:</strong> dependencies, upstreams, downstreams</p>
</li>
<li><p><strong>Infrastructure metadata:</strong> Kubernetes, databases, connector configuration</p>
</li>
</ul>
<p>Past incidents is the one worth calling out. Most tooling tells you what's broken now. DrDroid tells you this exact problem happened on March 14th and here's what resolved it. That history is the difference between an agent that hands you telemetry and one that hands you the answer.</p>
<h2>The Architecture</h2>
<p>Lets understand that your IDE is the client. DrDroid is the server and they talk over SSE.</p>
<img src="https://cdn.hashnode.com/uploads/covers/6a0febdbd8e265f60d9348f4/8693cdec-1b9a-4a02-8ead-ae0d0c109841.png" alt="" style="display:block;margin:0 auto" />

<p>The server pulls your production context into the editor. Nothing mutates through it. The agent reads your incident history. It does not get to restart your pods. For a tool sitting one query away from production, that boundary is the point.</p>
<p>Because it runs against your DrDroid instance, the same access you already have governs what the MCP server can see. It doesn't need any new surface or any separate permission to reason about. What it looks like in practice</p>
<p>You're in Claude Code with a traces open. You ask, in the same session:</p>
<blockquote>
<p>What happened with payment-service last night?</p>
</blockquote>
<p>The server returns, inline: the alerts that fired in the last 24 hours, the related deployments, the past incident investigations that match, and the service dependency graph. No alt-tab. No reassembly. The answer arrives next to the code you're about to change.</p>
<p>Each capability is a tool the agent invokes on its own during an investigation, or that you call directly.</p>
<p>The CI/CD set, for example:</p>
<pre><code class="language-plaintext">list_pipelines()              # discover CI/CD pipelines
get_pipeline_runs()           # recent executions
get_run_logs()                # detailed logs from a run
get_deployments()             # deployment status
get_pipeline_health_summary() # reliability trends over time
</code></pre>
<h2>Conclusion</h2>
<p>If you're a using Drdroid, point Claude Code at your Drdroid <code>/mcp</code> endpoint and the knowledge graph you've been building is live in your editor or you can build your own skills. Open a claude code session, ask it about last night's incident, and watch it pull the alerts, the deploys, and the matching past incident without you touching another tab.</p>
<p>The next time something breaks at 2 AM, you will not be looking for the context. It will already be where your cursor is.</p>
]]></content:encoded></item><item><title><![CDATA[Context Engine: How DrDroid's AI Agent leverages the Continuously Improving Knowledge Graph]]></title><description><![CDATA[Executive Summary
“Production systems” is a not a single component.
Kubernetes is a tangibly identifiable component with deterministic ways of operation. So is a database. And so is code. But “Product]]></description><link>https://notes.drdroid.io/context-engine-how-drdroid-s-ai-agent-leverages-the-continuously-improving-knowledge-graph</link><guid isPermaLink="true">https://notes.drdroid.io/context-engine-how-drdroid-s-ai-agent-leverages-the-continuously-improving-knowledge-graph</guid><category><![CDATA[#AIOps]]></category><category><![CDATA[AISRE]]></category><category><![CDATA[observability]]></category><dc:creator><![CDATA[Siddarth Jain]]></dc:creator><pubDate>Tue, 21 Apr 2026 09:02:21 GMT</pubDate><enclosure url="https://cdn.hashnode.com/uploads/covers/63200bf16c86d75accc7fd61/d1a77a36-2eb6-4ff9-aed9-aeb0a8b11081.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h3>Executive Summary</h3>
<p>“Production systems” is a not a single component.</p>
<p>Kubernetes is a tangibly identifiable component with deterministic ways of operation. So is a database. And so is code. But “Production Systems” is a complex orchestration of different and discrete components that interact with each other to produce desired results for a business or user.</p>
<p>Joining together pieces, some of which are continuously changing, means that mathematically, the risk of failure is non-linearly higher than the risk of the individual components.</p>
<p>To tackle that, it is quintessential to design, iterate and architect systems with resilience, fallbacks, reliability and visibility.</p>
<p>DrDroid helps engineers leverage AI effectively for automating and accelerating operations. To enable high quality results, DrDroid has a stateful knowledge layer (context engine) about different components of production systems for agents.</p>
<h3>The need for context engine</h3>
<p>Engineers are using DrDroid to debug production alerts and get sharp data backed RCAs. To ensure that the agent is able to run complex investigations, is grounded in data and is context aware, we realised that a lot of reality is hidden in the day-to-day changes in the system. Here a few examples of typical changes, in sequence of most obvious to less obvious:</p>
<ul>
<li><p><strong>Code Changes:</strong> Regular changes in a service in accordance with deliverables and goals</p>
</li>
<li><p><strong>Resource changes:</strong> A database table gets it’s memory upto the limit</p>
</li>
<li><p><strong>Indirect code changes:</strong> A third party dependency being used downstream in a service changes their API payload</p>
</li>
<li><p><strong>Indirect resource changes:</strong> A certificate expires on one of the machines where a critical service is running</p>
</li>
<li><p><strong>User change:</strong> A different user triggers an edge case in a product journey that never got used until now, leading to a failure for the user</p>
</li>
</ul>
<p>None of these issues are unheard of, or shockingly new to an engineer. The problem is that the impact could be visible at a very different point in the system compared to where the issue originated.</p>
<p>In fact, for very different reasons, the same outcome/impact/alert might get triggered at different times.</p>
<p>Being able to mitigate and avoid such issues often require more information than just the alert. <strong>This debugging is often a combination of system knowledge, deductive reasoning, tribal knowledge, specialised skills and real-time situational context.</strong></p>
<p>Added to that, engineers are expected to do this in high pressure situations or unexpected times as there are often business guarantees associated with these issues.</p>
<h3>Components within a context engine</h3>
<p>Our context engine is designed to proactively process and maintain an update knowledge layer about a company.</p>
<p>Here are some of the examples of what the context engine comprises of:</p>
<ol>
<li><p><strong>System Knowledge:</strong></p>
<img src="https://cdn.hashnode.com/uploads/covers/63200bf16c86d75accc7fd61/f5f32c6b-d478-4874-9062-47a953156a00.png" alt="" style="display:block;margin:0 auto" />

<ol>
<li><p><strong>Service Catalog:</strong> A list of every service, it's related teams and repositories, monitoring context and dependencies. This can auto-generated by fetching data across multiple tools like an APM or ArgoCD or Github repositories or Kubernetes. You can further edit the details if needed.</p>
<img src="https://cdn.hashnode.com/uploads/covers/63200bf16c86d75accc7fd61/e451f805-8c87-4513-b6d4-08ab2c8314cc.png" alt="" style="display:block;margin:0 auto" />
</li>
<li><p><strong>Infrastructure inventory:</strong> A list of all the resources, from Databases to VMs to kubernetes deployments and namespaces.</p>
</li>
<li><p><strong>Pre-instrumented metrics knowledge:</strong> The context engine is aware of all the metrics that are instrumented, as well as the dashboards and it's panels. This ensures that agent does not have to query prometheus by first principle but can re-use your existing dashboards too!</p>
</li>
<li><p><strong>Codebase context:</strong> The context engine continuously maintains knowledge of the capabilities and features within each repo, their structure. It additionally also stores correlations with other repositories (depending on the information available). Read more on <a href="https://docs.drdroid.io/policies/code-security#source-code-security">source code security</a>.</p>
</li>
<li><p><strong>Third party dependencies:</strong> Within DrDroid, you can maintain a list of all 3rd parties that your application depends upon.</p>
<img src="https://cdn.hashnode.com/uploads/covers/63200bf16c86d75accc7fd61/4348d90c-30c2-4644-affc-3329d03a4539.png" alt="" style="display:block;margin:0 auto" /></li>
</ol>
</li>
<li><p><strong>Assistance for deductive reasoning:</strong></p>
<ol>
<li><strong>Service &amp; infrastructure correlations:</strong> By leveraging the context of your inventory across infrastructure &amp; services, and your traces/service maps, correlations between services and infrastructure is auto-processed. This can be built using your existing service maps from APM / traces or using network mappers in your kubernetes infrastructure (available with our <a href="https://github.com/DrDroidLab/drd-vpc-agent">reverse proxy service</a>)</li>
</ol>
</li>
</ol>
<p><a class="embed-card" href="https://www.youtube.com/watch?v=Bt-G0ZRaj-8&amp;t=2s">https://www.youtube.com/watch?v=Bt-G0ZRaj-8&amp;t=2s</a></p>

<pre><code>2.  **Pre-compiled log patterns and fingerprints:** Your log patterns are overtime stored within the context engine, ensuring that the agent can re-use the past experience of log analysis to improve the future log search queries.
    
3.  **Pre-mapped data flows and critical user journeys:** The agent maintains context about the product/architecture and can even keep context of critical workflows across repositories. Additionally, these can be edited to add additional insights and knowledge.
    
</code></pre>
<ol>
<li><p><strong>Tribal Knowledge:</strong></p>
<ol>
<li><p><strong>Pre-existing runbooks and SOPs:</strong> All your existing runbooks and SOPs can be added to the context engine, either via syncing it with your wiki or uploading/creating documents in the account.</p>
</li>
<li><p><strong>Past Incidents &amp; investigations:</strong> The context engine is continuously updated as the agent runs alert investigations and RCAs. Insights of correlations and issues identified are stored for easy referencing and searching for the agent in the future.</p>
</li>
<li><p><strong>Alert Patterns:</strong> The context engine stores alerts from Slack channels and webhooks, ensuring the agent can see the pattern of a specific alert and differentiate an anomaly from a regular alert.</p>
</li>
<li><p><strong>Communication rosters:</strong> The context engine stores information about the teams and the users within the teams. Users can be auto-synced from Slack where as the teams can be simply configured in the platform.</p>
</li>
</ol>
</li>
<li><p><strong>Specialised skills</strong></p>
<ol>
<li><p><strong>Internal tools specific knowledge:</strong> You can store guidelines about how your internal CLI / tools operate. This would enable the agent to use the tools effectively, ensuring your production workflows can be replicated here.</p>
</li>
<li><p><strong>Best practices on trace analysis:</strong> Analysing traces can often be taxing for any engineer and require deep focus to identify the anomaly in them, across the time and spans. The context engine has specialised tools for trace analysis, enabling it to get insights on large volume of spans in a short duration with reduced token consumption.</p>
</li>
<li><p><strong>Observability tools usage:</strong> Every tool has it's own data structures and entity design. Best practices around using tools like Signoz, NewRelic, AWS, Kubernetes, etc. are pre-baked into the context engine.</p>
</li>
</ol>
</li>
<li><p><strong>Real-time situational context</strong></p>
<img src="https://cdn.hashnode.com/uploads/covers/63200bf16c86d75accc7fd61/96951e4c-1679-4ade-8f77-072cea7ba68f.png" alt="" style="display:block;margin:0 auto" />

<ol>
<li><p><strong>Code Changes:</strong> The context engine continuously processes PRs, commits and merges across repositories. This ensures that the agent can get a quick glance on recent changes.</p>
</li>
<li><p><strong>Infrastructure Changes:</strong> Addition of inventory in the infrastructure, from a database in your cloud to a namespace in your kubernetes cluster, the diff is continuously tracked.</p>
</li>
<li><p><strong>Ongoing issues and alerts:</strong> Alerts are processed and saved in real-time for the agent to use in real-time.</p>
</li>
<li><p><strong>Vendor downtime events:</strong> The context engine is continuously querying the status pages of the third party vendors identified in the platform. This ensures that the agent can be notified upfront during an investigation if a downstream vendor API is impacted.</p>
</li>
</ol>
</li>
</ol>
<h3>Where is the context engine used?</h3>
<p>How does the context engine get used in the platform:</p>
<ol>
<li><p><strong>Investigation Agent:</strong> The investigation agent is our flagship agent which can help you investigate any issue in your production system - from alert investigation &amp; service degradation to security review &amp; cost analysis.</p>
</li>
<li><p><strong>Alert Classification/Triaging Agent:</strong> Just send all your alerts to DrDroid. When you receive a large stream of alerts, this agent can classify your alert to be a noisy alert (and suppress it) or let that alert come through to you. It grounds every recommendation in actual system data. As you get started, it gives benefit of doubt to an alert being critical (so that a critical issue isn't missed) but over time, it learns from already triaged alerts, data queried and user feedback received. Users can also add specific "notes" at alert type level so that agent doesn't have to think from first principles when it gets started.</p>
</li>
<li><p><strong>Alert Grouping Agent:</strong> If you have 100+ relevant alerts coming in, rarely are they 100 unique issues. More often than not, these might end up becoming 7 or 19 issues. The Grouping agent listens to your alerts and tries to group it into actionable buckets - it often does it by root cause (found by investigation agent) or by the impacted component / where it fits in the topology. This could ensure that during an incident or when you get a surged stream of alerts where <em><strong>most</strong></em> of them are due to the same issue, the one other signal doesn't get missed.</p>
</li>
<li><p><strong>Continuous Improvement Agent:</strong> The core principles behind engineering operations is to react, learn and improve. Some teams often get too occupied with firefighting and do not get the bandwidth to work on proactive improvements of the system. This agent gives suggestions for improving your alerts, missing observability data for any of your new services, gives cost or security posture recommendations.</p>
</li>
<li><p><strong>Meta agent:</strong> This agent helps you to improve context on the DrDroid platform based on ongoing activity, which will in-turn, make all the other agents more powerful.</p>
</li>
</ol>
<h3>How is the context engine implemented?</h3>
<p>Read more about the design of the context engine and it's architecture in <a href="https://drdroid.io/dynamic-memory-retrieval">this blog</a>.</p>
]]></content:encoded></item><item><title><![CDATA[How DrDroid AI SRE Agent is specialised for Production Incidents & On-call Investigations]]></title><description><![CDATA[By working with 100s of engineers and their debugging problems, we iterated over DrDroid. The investigation agent assists engineers with complex analysis which are critical, time sensitive and have li]]></description><link>https://notes.drdroid.io/how-drdroid-ai-sre-agent-is-specialised-for-production-incidents-on-call-investigations</link><guid isPermaLink="true">https://notes.drdroid.io/how-drdroid-ai-sre-agent-is-specialised-for-production-incidents-on-call-investigations</guid><category><![CDATA[ai-sre]]></category><category><![CDATA[#AIOps]]></category><category><![CDATA[observability]]></category><category><![CDATA[monitoring]]></category><category><![CDATA[logging]]></category><dc:creator><![CDATA[Siddarth Jain]]></dc:creator><pubDate>Thu, 09 Apr 2026 12:32:00 GMT</pubDate><enclosure url="https://cdn.hashnode.com/uploads/covers/63200bf16c86d75accc7fd61/4502861e-8549-49d8-a504-3e26c79cc16a.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>By working with 100s of engineers and their debugging problems, we iterated over DrDroid. The investigation agent assists engineers with complex analysis which are critical, time sensitive and have little room for error. Here are some of the things that help the investigation agent perform well:</p>
<h2><strong>1. Specialized Debugging Tools &amp; Skills</strong></h2>
<p>Production incidents often require analyzing large volumes of logs, traces, and metrics. This can be token-intensive and time-consuming. Most LLMs and agentic frameworks hit context window limits quickly, can't process production-scale data, and lose quality when analyzing large datasets.</p>
<p><strong>How DrDroid Solves This:</strong></p>
<h4><strong>Pre-Built Aggregate Analysis Tools</strong></h4>
<p>Instead of feeding raw logs into an LLM, DrDroid has tools designed specifically for handling large-volume logs with:</p>
<ul>
<li><p>Built-in log aggregation and pattern detection</p>
</li>
<li><p>Complex trace analysis across distributed systems</p>
</li>
<li><p>Large-volume metrics analysis with outlier detection and ML techniques</p>
</li>
</ul>
<h4><strong>Specialized Investigation Skills</strong></h4>
<p>Our agent has domain-specific skills built from working with hundreds of engineers on real debugging problems:</p>
<ul>
<li><p>How to query and analyze traces in Signoz or Datadog</p>
</li>
<li><p>How to navigate APM data efficiently</p>
</li>
<li><p>How to correlate metrics across multiple monitoring tools</p>
</li>
</ul>
<p><strong>Real Impact:</strong> The agent can process 100,000+ log lines in seconds and surface the 5 relevant errors—something that would exhaust a generic LLM's context window.</p>
<h2>2. Code &amp; Application Awareness</h2>
<p>Most LLMs and agents start every investigation from zero, with only the context of the prompt and sometimes markdown files.</p>
<h3><strong>How DrDroid Solves This:</strong></h3>
<h4>Automatic Code Context Generation</h4>
<p>Even before your first chat with the agent, DrDroid builds knowledge of:</p>
<ul>
<li><p>What each repository does</p>
</li>
<li><p>What capabilities, APIs, features, and workflows each repo covers</p>
</li>
<li><p>Programming languages, frameworks, and file structures</p>
</li>
<li><p>Connections between multiple repositories (discovered via traces and logs)</p>
</li>
</ul>
<h4>Business Workflow Understanding</h4>
<p>You can ask DrDroid to build context around critical business and product workflows. The agent understands:</p>
<ul>
<li><p>"The checkout flow involves payment-service, inventory-service, and notification-service"</p>
</li>
<li><p>"When users report 'payment stuck,' check these three services in this order"</p>
</li>
</ul>
<p><strong>Real Impact:</strong> When an alert fires on "payment-service," DrDroid already knows what that service does, which other services depend on it, and where to look for root causes.</p>
<h2>3. Infrastructure &amp; Resource Awareness</h2>
<p>An LLM with MCP connections doesn't know which apps run in which Kubernetes clusters, which databases are in which cloud providers, or how your infrastructure is organised. It needs to query multiple tools (costing time and token) and explore it before it's able to answer that.</p>
<p><strong>How DrDroid Solves This:</strong></p>
<h4>Auto-Discovery of Infrastructure</h4>
<p>DrDroid continuously maps:</p>
<ul>
<li><p>Apps hosted in different Kubernetes clusters</p>
</li>
<li><p>Databases and their cloud providers</p>
</li>
<li><p>Service dependencies and communication patterns</p>
</li>
<li><p>Network topology and resource relationships</p>
</li>
</ul>
<h4>Service Map &amp; Dependency Graph</h4>
<p>The agent can answer questions like:</p>
<ul>
<li><p>"Which services depend on the payments database?"</p>
</li>
<li><p>"If eu-west-1 goes down, what's affected?"</p>
</li>
<li><p>"Show me all services running in the production cluster"</p>
</li>
</ul>
<p><strong>Real Impact:</strong> During an incident, the agent instantly <strong>knows the blast radius</strong> and <strong>which downstream services might be affected</strong>—without you having to explain your architecture.</p>
<h2>4. Past Alert &amp; Incident Pattern Recognition</h2>
<p>With generic agents, every investigation is independent. It has no memory of past incidents or patterns.</p>
<p><strong>How DrDroid Solves This:</strong> Searchable Alert History</p>
<p>The agent has access to:</p>
<ul>
<li><p>All alerts since platform enablement</p>
</li>
<li><p>Past incidents and their resolutions</p>
</li>
<li><p>RCAs and postmortems (from Confluence, docs, or previous investigations)</p>
</li>
<li><p>Understanding patterns in alerts</p>
</li>
</ul>
<p>When similar issues occur, the agent can say:</p>
<ul>
<li><p>"This looks similar to the incident from Jan 15th where the Redis cache was full"</p>
</li>
<li><p>"Last time this alert fired, the root cause was a config change in service X"</p>
</li>
</ul>
<p><strong>Real Impact:</strong> Repeat incidents get resolved faster because the agent learns from past investigations.</p>
<h2>5. Continually Learning System</h2>
<p>DrDroid improves with every investigation.</p>
<h4>Active Learning from Your Environment</h4>
<p>The agent continuously creates notes and memory from:</p>
<ul>
<li><p>Recent commits and merges in your applications</p>
</li>
<li><p>Investigations and conversations with the agent</p>
</li>
<li><p>Human conversations in Slack channels (optional)</p>
</li>
</ul>
<h4>Contextual Memory Storage</h4>
<p>Everything is stored with metadata:</p>
<ul>
<li><p>Timestamp</p>
</li>
<li><p>Related entities (services, databases, clusters)</p>
</li>
<li><p>Related team and people</p>
</li>
<li><p>Relevant tags and categories</p>
</li>
</ul>
<p><strong>Real Impact:</strong> The agent gets smarter every week. After a month, it knows your environment better than most new engineers.</p>
<h2>6. Context Compaction (1M+ Token Conversations)</h2>
<p>Typically, agents do the following with the problem of large context windows:</p>
<ul>
<li><p>Summarize the entire conversation (losing critical context) or</p>
</li>
<li><p>Hit token limits and can't continue</p>
</li>
<li><p>Slow down dramatically as conversations grow</p>
</li>
</ul>
<p>With production telemetry data, this can often happen.</p>
<p><strong>How DrDroid Solves This:</strong></p>
<h4>Intelligent Compression Without Context Loss</h4>
<ul>
<li><p>Tool calls are compressed (only IDs and summaries preserved)</p>
</li>
<li><p>Reasoning and train of thought remain intact (no summarization)</p>
</li>
<li><p>Agent maintains full context even beyond 1M tokens</p>
</li>
<li><p>Smart Tool-Level Compaction</p>
</li>
</ul>
<p>Our tools have built-in context management:</p>
<ul>
<li><p>Logging tool has grep/search capability over large volumes</p>
</li>
<li><p>Agent can "eyeball and search" logs instead of loading everything into context</p>
</li>
</ul>
<p><strong>Real Impact:</strong> You can have a 2-hour debugging session with 500+ tool calls, and the agent never loses context or slows down.</p>
<h2>7. Multi-Channel Conversations with Shareability</h2>
<p>The agent is designed to work from your place of convenience:</p>
<ul>
<li><p>Slack DMs</p>
</li>
<li><p>Thread replies to alerts</p>
</li>
<li><p>Web UI</p>
</li>
<li><p>CLI (coming soon)</p>
</li>
<li><p>API triggers</p>
</li>
<li><p>Voice calls (coming soon)</p>
</li>
</ul>
<p><strong>Seamless Sharing</strong></p>
<p>Any investigation can be:</p>
<ul>
<li><p>Shared with teammates for review</p>
</li>
<li><p>Linked in postmortems</p>
</li>
<li><p>Referenced in future incidents</p>
</li>
</ul>
<p><strong>Real Impact:</strong> When someone gets paged, they can see the auto-investigation that already ran in the Slack thread—no need to DM the agent separately.</p>
<h2>8. Automated Investigations</h2>
<p>DrDroid can run proactively or via automated triggers enabling proactive visibility for your team:</p>
<ul>
<li><p>Alert fires in PagerDuty/OpsGenie → Investigation starts automatically</p>
</li>
<li><p>Cron-based health checks → Agent investigates on schedule</p>
</li>
<li><p>Custom triggers via API or webhooks</p>
</li>
</ul>
<p><strong>Real Impact:</strong> Agent can detect issues even without alerts; By the time you open the alert, the agent has already investigated and summarized the likely root cause.</p>
<h2>9. Smart Model Switching (85% Cost Savings)</h2>
<p>LLMs have been commoditised and the SOTA model is not necessarily required for every investigation. DrDroid smartly chooses between different LLMs based on investigation complexity</p>
<ul>
<li><p>Simple tasks → Faster, cheaper models</p>
</li>
<li><p>Complex reasoning → State-of-the-art models</p>
</li>
</ul>
<p><strong>Real Impact:</strong> Up to 85% token savings compared to always using frontier models, with no degradation in investigation quality.</p>
<h2>10. Dedicated File System &amp; Memory</h2>
<p>Memory management for a large scale infrastructure requires a structured approach.</p>
<p>DrDroid has a Persistent Knowledge Base - All context, memory, investigations, and alerts are stored and accessible:</p>
<ul>
<li><p>Agent can navigate past investigations like files</p>
</li>
<li><p>Search across all historical data</p>
</li>
<li><p>Reference previous findings instantly</p>
</li>
</ul>
<p><strong>Real Impact:</strong> "Show me all investigations related to database timeouts in the last 30 days" returns instant results.</p>
<h2>11. Coding Sub-Agent for Hotfixes</h2>
<p>Coding agent operates very different from a production investigation agent. DrDroid comes pre-packaged with a coding agent connected to the investigation agent.</p>
<p><strong>Real-Time Coding Agent When needed:</strong></p>
<ul>
<li><p>Spins up a coding agent in an ephemeral sandbox</p>
</li>
<li><p>Reviews the full repository</p>
</li>
<li><p>Creates hotfix PRs with proper context</p>
</li>
</ul>
<p><strong>Real Impact:</strong> During an incident, the agent can say "I found the bug in payment-service line 247—here's a PR to fix it."</p>
<h2>12. Remote Machine &amp; Kubernetes Access</h2>
<p>Often, data needs to be reviewed on a remote machine or kubernetes cluster for production incidents. These might be inaccessible or sensitive.</p>
<p><strong>How DrDroid Solves This:</strong> Direct Infrastructure Access without token access to the agent</p>
<ul>
<li><p>Execute commands on remote machines via SSH (keys are not exposed to the agent)</p>
</li>
<li><p>Query read-only Kubernetes clusters directly</p>
</li>
<li><p>Access VMs and clusters within your VPC via reverse proxy</p>
</li>
</ul>
<p><strong>Real Impact:</strong> "Check disk space on prod-api-01" → Agent SSHs in, runs the command, and returns results. No manual execution needed.</p>
<h2>13. Image Support for Dashboard Analysis</h2>
<p>You might want to debug an issue with a screenshot shared by the customer as the starting point.</p>
<p><strong>How DrDroid Solves This:</strong> DrDroid agent support image processing from Slack or UI.</p>
<ul>
<li><p>Your product showing an error</p>
</li>
<li><p>A Grafana dashboard</p>
</li>
<li><p>A monitoring alert</p>
</li>
</ul>
<p>The agent analyses it and continues the investigation from there.</p>
<p><strong>Real Impact:</strong> "Here's what the user is seeing" → Agent understands the UI issue and investigates the backend cause.</p>
<h2>14. Granular Access Control &amp; RBAC</h2>
<p>Production systems debugging come with sensitive data and access management. DrDroid ensures only the right people have the right access while debugging.</p>
<p><strong>How DrDroid Solves This:</strong></p>
<ul>
<li><p><strong>Read commands:</strong> Execute without approval (safe exploration)</p>
</li>
<li><p><strong>Write commands:</strong> Require RBAC approval per your policy</p>
</li>
<li><p><strong>SSO integration:</strong> Syncs with your internal permissions</p>
</li>
<li><p><strong>Audit logs:</strong> Track who did what</p>
</li>
</ul>
<p><strong>Real Impact:</strong> Junior engineers can investigate safely, while dangerous operations require senior approval.</p>
<h2>15. Third-Party Vendor Status Tracking</h2>
<p>Often production incidents can be partly caused due to 3rd party downtimes or issues.</p>
<p><strong>How DrDroid Solves This:</strong> Connected to Vendor Statuspages</p>
<ul>
<li><p>Tracks 150+ status pages for your third-party vendors: Stripe, AWS, Datadog, MongoDB Atlas, etc.</p>
</li>
<li><p>Flags when vendor issues might be causing downstream impact</p>
</li>
</ul>
<p><strong>Real Impact:</strong> "Is this our issue or Stripe's?" → Agent checks Stripe's status page and correlates timing.</p>
<h2>16. Automated Quality Evaluation</h2>
<p>Production Agents need to come with quality guarantees for the team to track and trust.</p>
<p><strong>How DrDroid Solves This:</strong> LLM-Based Evals on Every Investigation</p>
<p>Every investigation is automatically evaluated for:</p>
<ul>
<li><p>Accuracy Safety Errors or hallucinations</p>
</li>
<li><p>Central teams get visibility into investigation quality and improvement opportunities.</p>
</li>
</ul>
<p><strong>Real Impact:</strong> Platform team can see "Investigation quality is 94% this month, down from 97% last month—let's review the low-scoring investigations."</p>
<h2>17. User Feedback &amp; Team Visibility</h2>
<p>Context within DrDroid can continuously improve over time. But for that to improve, tracking and acting upon user feedback is critical.</p>
<p><strong>How DrDroid Solves This:</strong> Collaborative Quality Control</p>
<ul>
<li><p>Every investigation can be upvoted/downvoted</p>
</li>
<li><p>Feedback back to central team helps improve agent context</p>
</li>
</ul>
<p><strong>Real Impact:</strong> Central team &amp; managers have visibility on confidence and impact of AI on the engineers.</p>
<h2>18. Reasoning Lifecycle &amp; Audit Trail</h2>
<p>Production incident investigations cannot be led to be incorrect due to "hallucinations" or "guesses" by an LLM. DrDroid ensures that every reasoning and logic by the LLM is grounded in facts and data.</p>
<p><strong>How DrDroid Solves This:</strong> Transparent Investigation Path</p>
<p>The agent tracks:</p>
<ul>
<li><p>What data it queried</p>
</li>
<li><p>Why each data point was relevant</p>
</li>
<li><p>What hypothesis it built from each finding</p>
</li>
<li><p>How it reached its conclusion</p>
</li>
</ul>
<p><strong>Real Impact:</strong> You can backtrack through the investigation to validate correctness, spot gaps, or understand the agent's reasoning.</p>
<p><strong>Summary:</strong> Why DrDroid is Purpose-Built for Production</p>
<table>
<thead>
<tr>
<th>Capability</th>
<th>DrDroid Investigation Agent</th>
</tr>
</thead>
<tbody><tr>
<td>Code awareness</td>
<td>Auto-discovers repos, APIs, dependencies</td>
</tr>
<tr>
<td>Infrastructure knowledge</td>
<td>Knows your K8s, cloud, databases</td>
</tr>
<tr>
<td>Log/metric analysis</td>
<td>Specialized tools for production-scale data</td>
</tr>
<tr>
<td>Memory of past incidents</td>
<td>Full history + pattern learning</td>
</tr>
<tr>
<td>Context window</td>
<td>1M+ tokens with intelligent compaction</td>
</tr>
<tr>
<td>Cost optimization</td>
<td>Smart switching, 85% savings</td>
</tr>
<tr>
<td>Permissions &amp; RBAC</td>
<td>Enterprise-grade access control</td>
</tr>
<tr>
<td>Auto-triggered investigations</td>
<td>Yes—from alerts, cron, API</td>
</tr>
<tr>
<td>Quality control</td>
<td>Automated evals + team feedback</td>
</tr>
<tr>
<td>Infrastructure execution</td>
<td>Direct SSH, K8s, API access</td>
</tr>
</tbody></table>
<p>What This Means for Your Team:</p>
<ul>
<li><p>Every investigation starts with past context</p>
</li>
<li><p>You do not have to guide the LLM or explain your architecture every time</p>
</li>
<li><p>It can handle production-scale logs or metrics</p>
</li>
<li><p>It supports automation or proactive help</p>
</li>
<li><p>It works across different channels where your team lives</p>
</li>
</ul>
<p>Ready to See the agent?</p>
<p>DrDroid is a purpose-built investigation agent that understands your infrastructure, learns from your incidents, and gets smarter every day.</p>
<h2>Next steps:</h2>
<ul>
<li><p><a href="https://www.youtube.com/@DrDroidDev">Watch platform demo videos</a></p>
</li>
<li><p>Check our <a href="https://drdroid.io/integrations">MCP Servers &amp; integrations</a></p>
</li>
<li><p>Read the <a href="https://docs.drdroid.io/">documentation</a></p>
</li>
<li><p>See customer <a href="https://drdroid.io/case-studies">case studies</a></p>
</li>
</ul>
<p>Want to see how it works with your stack?</p>
<p>Setup &amp; go live takes 1-2 hours for smaller teams, &lt; 1 week for enterprises. Get started <a href="https://drdroid.io/">here</a>.</p>
]]></content:encoded></item><item><title><![CDATA[DrDroid: How AI SRE Helps Engineers who are on-call for production monitoring]]></title><description><![CDATA[Day 0 — Incident Response & Firefighting
Something is broken right now. You need answers fast.



#
Your Task
How Doctor Droid Helps You



1
"Our API latency spiked, what changed?"
Ask the agent — it]]></description><link>https://notes.drdroid.io/drdroid-how-ai-sre-helps-engineers-who-are-on-call-for-production-monitoring</link><guid isPermaLink="true">https://notes.drdroid.io/drdroid-how-ai-sre-helps-engineers-who-are-on-call-for-production-monitoring</guid><category><![CDATA[Open Source]]></category><category><![CDATA[monitoring]]></category><category><![CDATA[observability]]></category><category><![CDATA[#AIOps]]></category><category><![CDATA[ai-sre]]></category><category><![CDATA[ai agents]]></category><dc:creator><![CDATA[Siddarth Jain]]></dc:creator><pubDate>Tue, 07 Apr 2026 10:40:52 GMT</pubDate><enclosure url="https://cdn.hashnode.com/uploads/covers/63200bf16c86d75accc7fd61/4040cb32-0484-4703-b287-0f6dceca395d.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h2>Day 0 — Incident Response &amp; Firefighting</h2>
<p><em>Something is broken right now. You need answers fast.</em></p>
<table>
<thead>
<tr>
<th><strong>#</strong></th>
<th><strong>Your Task</strong></th>
<th><strong>How Doctor Droid Helps You</strong></th>
</tr>
</thead>
<tbody><tr>
<td>1</td>
<td><strong>"Our API latency spiked, what changed?"</strong></td>
<td>Ask the agent — it pulls recent deployments from ArgoCD/GitHub Actions, checks Grafana/Datadog metrics, and shows you what changed around the time latency spiked. No need to open 5 tabs.</td>
</tr>
<tr>
<td>2</td>
<td><strong>"Which pods are crashing in production?"</strong></td>
<td>Ask the agent — it lists failing pods, pulls their logs, shows recent K8s events, and surfaces restart counts. You get a full picture in one response.</td>
</tr>
<tr>
<td>3</td>
<td><strong>"Is the database the bottleneck?"</strong></td>
<td>Ask the agent — it runs slow query analysis on your Postgres/MySQL, checks connection pool usage, and correlates with application error rates from Datadog or New Relic.</td>
</tr>
<tr>
<td>4</td>
<td><strong>"I got paged, what's actually going on?"</strong></td>
<td>The agent auto-investigates when an alert fires. By the time you open it, there's already a summary with metrics, logs, and likely root cause pulled from your connected sources.</td>
</tr>
<tr>
<td>5</td>
<td><strong>"Are other services affected too?"</strong></td>
<td>Ask the agent — it checks health across your connected sources (Grafana, CloudWatch, K8s, Datadog) and tells you which services are degraded vs healthy.</td>
</tr>
<tr>
<td>6</td>
<td><strong>"I need to check CloudWatch logs for this error"</strong></td>
<td>Tell the agent the error pattern and time range — it queries CloudWatch Logs, Loki, or Elasticsearch directly and returns matching entries. No console login needed.</td>
</tr>
<tr>
<td>7</td>
<td><strong>"Run this PromQL/NRQL query for me"</strong></td>
<td>Give the agent your query — it executes against Prometheus, Grafana, New Relic, or Datadog and returns results inline. Great for quick checks during incidents.</td>
</tr>
<tr>
<td>8</td>
<td><strong>"SSH into the box and check disk space"</strong></td>
<td>Tell the agent the command — it runs Bash commands on remote hosts via SSH and returns output. No need to find the SSH key or remember the hostname. Works even when you don’t have laptop access.</td>
</tr>
<tr>
<td>9</td>
<td><strong>"Notify the team on Slack about this outage"</strong></td>
<td>Tell the agent what to post and where — it sends a formatted message to your Slack channel with the context you provide.</td>
</tr>
<tr>
<td>10</td>
<td><strong>"Escalate this to PagerDuty"</strong></td>
<td>The agent creates or escalates a PagerDuty/OpsGenie incident based on the investigation findings. You don't need to context-switch to the PagerDuty UI.</td>
</tr>
<tr>
<td>11</td>
<td><strong>"Create a JIRA ticket for the post-mortem"</strong></td>
<td>Tell the agent the summary — it creates a JIRA ticket with the incident details, investigation findings, and relevant links.</td>
</tr>
<tr>
<td>12</td>
<td><strong>"Did the last deploy cause this?"</strong></td>
<td>Ask the agent — it checks the latest GitHub PR merges, Jenkins builds, ArgoCD sync status, and correlates timestamps with when the issue started.</td>
</tr>
<tr>
<td>13</td>
<td><strong>"Roll back the deployment"</strong></td>
<td>Tell the agent to trigger a rollback pipeline — it kicks off the Jenkins build or GitHub Actions workflow you specify.</td>
</tr>
<tr>
<td>14</td>
<td><strong>"Check if this API endpoint is responding"</strong></td>
<td>Give the agent the URL — it makes an HTTP call and tells you the status code, response time, and body. Works for any internal API.</td>
</tr>
</tbody></table>
<h2>Day 1 — Operational Tasks &amp; Maintenance</h2>
<p><em>No fire, but you need to keep things running smoothly.</em></p>
<table>
<thead>
<tr>
<th><strong>#</strong></th>
<th><strong>Your Task</strong></th>
<th><strong>How Doctor Droid Helps You</strong></th>
</tr>
</thead>
<tbody><tr>
<td>1</td>
<td><strong>"What's the current state of our K8s cluster?"</strong></td>
<td>Ask the agent — it shows pod status across namespaces, node resource usage, recent events, and any pods in CrashLoopBackOff or Pending state.</td>
</tr>
<tr>
<td>2</td>
<td><strong>"Are all our data sources healthy?"</strong></td>
<td>The agent tests connectivity to every configured connector every 10 seconds. Ask it for the current status, or set up Slack alerts for failures.</td>
</tr>
<tr>
<td>3</td>
<td><strong>"Show me all Grafana dashboards we have"</strong></td>
<td>Ask the agent — it returns the full inventory of dashboards, datasources, and folders it auto-discovered. Same for Datadog monitors, K8s resources, DB schemas, etc.</td>
</tr>
<tr>
<td>4</td>
<td><strong>"How big are our database tables getting?"</strong></td>
<td>Ask the agent to run a table size query on your Postgres/MySQL/ClickHouse — it returns sizes, row counts, and index usage.</td>
</tr>
<tr>
<td>5</td>
<td><strong>"Any slow queries running right now?"</strong></td>
<td>Ask the agent — it checks pg_stat_activity or equivalent on your database and shows long-running queries with their duration and state.</td>
</tr>
<tr>
<td>6</td>
<td><strong>"Check if our Jenkins pipelines are green"</strong></td>
<td>Ask the agent — it pulls recent build status from Jenkins and tells you which jobs passed, failed, or are stuck.</td>
</tr>
<tr>
<td>7</td>
<td><strong>"Is our ArgoCD app in sync?"</strong></td>
<td>Ask the agent — it checks sync status and health for your ArgoCD applications. Flags any that are out-of-sync or degraded.</td>
</tr>
<tr>
<td>8</td>
<td><strong>"Pull logs from the payment service for the last hour"</strong></td>
<td>Tell the agent the service and time range — it queries Loki, CloudWatch, or Elasticsearch and returns the logs. No need to remember log group names.</td>
</tr>
<tr>
<td>9</td>
<td><strong>"What GitHub PRs were merged today?"</strong></td>
<td>Ask the agent — it queries your GitHub repos and lists merged PRs with authors, titles, and timestamps.</td>
</tr>
<tr>
<td>10</td>
<td><strong>"Trigger a build for the staging environment"</strong></td>
<td>Tell the agent which Jenkins job or GitHub Actions workflow to run — it triggers it and reports back the status.</td>
</tr>
<tr>
<td>11</td>
<td><strong>"Send a daily cluster health report to Slack"</strong></td>
<td>Set up a scheduled playbook — the agent runs K8s health checks daily and posts a summary to your Slack channel automatically.</td>
</tr>
<tr>
<td>12</td>
<td><strong>"Check MongoDB replica set health"</strong></td>
<td>Ask the agent — it runs the appropriate commands against your MongoDB instance and returns replica set status and lag.</td>
</tr>
<tr>
<td>13</td>
<td><strong>"List all CloudWatch alarms in ALARM state"</strong></td>
<td>Ask the agent — it queries CloudWatch and returns currently firing alarms with their metric details and thresholds.</td>
</tr>
<tr>
<td>14</td>
<td><strong>"What Datadog monitors are in alert?"</strong></td>
<td>Ask the agent — it checks your Datadog monitors and lists any that are currently alerting or warning.</td>
</tr>
<tr>
<td>15</td>
<td><strong>"Run this custom SQL query on production"</strong></td>
<td>Give the agent the query — it executes against your connected Postgres, MySQL, ClickHouse, or BigQuery and returns results as a table.</td>
</tr>
<tr>
<td>16</td>
<td><strong>"Check disk and memory on our VMs"</strong></td>
<td>Tell the agent to run df -h and free -m via Bash — it SSHs in and returns the output.</td>
</tr>
<tr>
<td>17</td>
<td><strong>"Update this JIRA ticket with today's progress"</strong></td>
<td>Tell the agent the ticket ID and comment — it adds the update to JIRA without you opening the browser.</td>
</tr>
<tr>
<td>18</td>
<td><strong>"What new resources were created in our K8s cluster this week?"</strong></td>
<td>Ask the agent — it compares the current asset inventory with the previous discovery and highlights new pods, services, or deployments.</td>
</tr>
</tbody></table>
<hr />
<h2>Day 2 — Automation, Optimization &amp; Reliability</h2>
<p><em>You want to stop doing repetitive things and build resilience.</em></p>
<table>
<thead>
<tr>
<th><strong>#</strong></th>
<th><strong>Your Task</strong></th>
<th><strong>How Doctor Droid Helps You</strong></th>
</tr>
</thead>
<tbody><tr>
<td>1</td>
<td><strong>"Automate our incident response runbook"</strong></td>
<td>Build a runbook as per your internal process: when alert fires → pull metrics from Grafana → check K8s pods → query DB → post findings to Slack. Runs automatically on every alert.</td>
</tr>
<tr>
<td>2</td>
<td><strong>"Auto-restart pods when OOM detected"</strong></td>
<td>Set up a workflow: if K8s event shows OOMKilled → restart the pod → notify on Slack → log to JIRA. No human in the loop.</td>
</tr>
<tr>
<td>3</td>
<td><strong>"Scale up when CPU crosses 80%"</strong></td>
<td>Create a conditional workflow: check CPU metrics from Prometheus/CloudWatch → if above threshold → trigger scaling via K8s or API call → confirm on Slack.</td>
</tr>
<tr>
<td>4</td>
<td><strong>"Track our SLOs across services"</strong></td>
<td>Set up scheduled playbooks that query Prometheus or Datadog for error rate and latency → calculate against SLO targets → post weekly reports to Slack.</td>
</tr>
<tr>
<td>5</td>
<td><strong>"Alert me before we hit our error budget"</strong></td>
<td>Build a workflow that checks error budget consumption daily → if &gt; 80% consumed → post warning to Slack and create a JIRA ticket.</td>
</tr>
<tr>
<td>6</td>
<td><strong>"Correlate deploys with performance changes"</strong></td>
<td>Set up a playbook that triggers after every deploy → compares pre/post metrics from Grafana/Datadog → flags regressions automatically.</td>
</tr>
<tr>
<td>7</td>
<td><strong>"Stop SSHing into boxes for the same checks"</strong></td>
<td>Convert your common SSH commands into playbooks — disk space, process checks, log tailing. Run them from Doctor Droid with one click or on schedule.</td>
</tr>
<tr>
<td>8</td>
<td><strong>"I keep opening 5 dashboards for the same investigation"</strong></td>
<td>Build a single playbook that queries all 5 sources and gives you a combined view. Next time, ask the agent instead of opening dashboards.</td>
</tr>
<tr>
<td>9</td>
<td><strong>"Validate infrastructure after every Terraform apply"</strong></td>
<td>Set up a post-deploy playbook: check K8s resources → verify CloudWatch alarms exist → test endpoints → report pass/fail.</td>
</tr>
<tr>
<td>10</td>
<td><strong>"Audit our K8s RBAC and network policies weekly"</strong></td>
<td>Schedule a playbook that lists RBAC bindings and network policies → compares against expected state → flags drift on Slack.</td>
</tr>
<tr>
<td>11</td>
<td><strong>"Auto-create a JIRA ticket when a deploy fails"</strong></td>
<td>Build a workflow: monitor Jenkins/GitHub Actions → if build fails → create JIRA ticket with build logs and assign to the team.</td>
</tr>
<tr>
<td>12</td>
<td><strong>"Map out which services talk to which"</strong></td>
<td>Enable Network Mapper — it discovers service-to-service communication in K8s and shows you the dependency graph.</td>
</tr>
<tr>
<td>13</td>
<td><strong>"Check if our Datadog monitors match our runbook"</strong></td>
<td>Ask the agent to list all Datadog monitors — compare with your documented expectations. Set this up as a weekly drift check.</td>
</tr>
<tr>
<td>14</td>
<td><strong>"Reduce alert fatigue for the team"</strong></td>
<td>Use alert grouping and conditional workflows — similar alerts get grouped, only actionable ones reach Slack/PagerDuty.</td>
</tr>
<tr>
<td>15</td>
<td><strong>"Automate capacity reporting for leadership"</strong></td>
<td>Schedule a monthly playbook: query CloudWatch/K8s for resource trends → query BigQuery for cost data → format and send via email or Slack.</td>
</tr>
<tr>
<td>16</td>
<td><strong>"Clear application cache when memory exceeds threshold"</strong></td>
<td>Build a workflow: check memory metrics → if above limit → make API call to flush cache → verify memory dropped → notify on Slack.</td>
</tr>
<tr>
<td>17</td>
<td><strong>"Run synthetic health checks every 5 minutes"</strong></td>
<td>Schedule a playbook that hits your critical endpoints via HTTP, checks response codes and latency, and alerts on Slack if anything degrades.</td>
</tr>
<tr>
<td>18</td>
<td><strong>"Onboard new services faster"</strong></td>
<td>When a new service deploys, the agent auto-discovers it in K8s, finds its Grafana dashboards and Datadog monitors, and catalogs everything. You see it in your inventory immediately.</td>
</tr>
</tbody></table>
<h2><strong>What's Next?</strong></h2>
<p>DrDroid helps your team move from <strong>reactive firefighting</strong> to <strong>proactive operations</strong>. Whether you're debugging an incident at 3 AM or building automation to prevent the next one, DrDroid becomes your team's operational co-pilot.</p>
<p><strong>Ready to see how DrDroid works with your stack?</strong></p>
<ul>
<li><p><a href="https://www.youtube.com/@DrDroidDev">Watch platform demo videos</a></p>
</li>
<li><p><a href="https://drdroid.io/integrations">Check our integrations</a></p>
</li>
<li><p><a href="https://docs.drdroid.io/">Read the documentation</a></p>
</li>
<li><p><a href="https://drdroid.io/case-studies">See customer case studies</a></p>
</li>
</ul>
<p><strong>Want to try it out?</strong> Setup takes 1-2 hours and you'll see value from your first investigation. <a href="https://drdroid.io/">Get started here</a>.</p>
]]></content:encoded></item><item><title><![CDATA[KubeCon + CloudNativeCon Europe 2026 Guide – Amsterdam
]]></title><description><![CDATA[Agenda Strategy, Tracks, Networking & SRE Playbook
UpdatedMarch 2026 • 6 min read

On this page

Doctor Droid’s Guide to KubeCon + CloudNativeCon Europe 2026

Overview

Event Schedule

Who Should Atte]]></description><link>https://notes.drdroid.io/kubecon-cloudnativecon-europe-2026-guide-amsterdam</link><guid isPermaLink="true">https://notes.drdroid.io/kubecon-cloudnativecon-europe-2026-guide-amsterdam</guid><category><![CDATA[kubeconeurope]]></category><category><![CDATA[Kubernetes]]></category><category><![CDATA[observability]]></category><dc:creator><![CDATA[Karan Sirohi]]></dc:creator><pubDate>Thu, 19 Mar 2026 06:59:28 GMT</pubDate><content:encoded><![CDATA[<p>Agenda Strategy, Tracks, Networking &amp; SRE Playbook</p>
<p><strong>Updated</strong><br />March 2026 • 6 min read</p>
<hr />
<h2>On this page</h2>
<ul>
<li><p>Doctor Droid’s Guide to KubeCon + CloudNativeCon Europe 2026</p>
</li>
<li><p>Overview</p>
</li>
<li><p>Event Schedule</p>
</li>
<li><p>Who Should Attend (and Who Can Skip)</p>
</li>
<li><p>How to Build Your Agenda (Without Overloading Yourself)</p>
</li>
<li><p>Track Strategy for SRE &amp; Platform Teams</p>
</li>
<li><p>Co-located Events: Where Specialists Get Leverage</p>
</li>
<li><p>Solutions Showcase: How to Avoid Vendor Fatigue</p>
</li>
<li><p>Networking Strategy for Engineering Leaders</p>
</li>
<li><p>Amsterdam Logistics Checklist</p>
</li>
<li><p>Post-Conference Execution Plan</p>
</li>
<li><p>Visit the Doctor Droid Booth</p>
</li>
</ul>
<hr />
<h2>Doctor Droid’s Guide to KubeCon + CloudNativeCon Europe 2026</h2>
<img src="https://cdn.hashnode.com/uploads/covers/66c6f53299c4280b93ee6b32/fda0dafb-b20e-4d4c-9e80-437d720ec783.png" alt="" style="display:block;margin:0 auto" />

<p>KubeCon + CloudNativeCon Europe is heading to <strong>Amsterdam, Netherlands, from 23–26 March 2026</strong>.</p>
<p>If you’re on <strong>platform engineering, SRE, DevOps, or cloud architecture teams</strong>, this event is still one of the best places to <strong>compress a year of learning into four days</strong>.</p>
<p>This guide is for <strong>practitioners who want outcomes, not conference FOMO</strong>:</p>
<ul>
<li><p>what to prioritize</p>
</li>
<li><p>how to choose tracks</p>
</li>
<li><p>how to plan networking</p>
</li>
<li><p>and how to convert conference notes into production improvements</p>
</li>
</ul>
<hr />
<h2>Overview</h2>
<p>KubeCon + CloudNativeCon is the <strong>flagship conference organized by the Cloud Native Computing Foundation (CNCF)</strong>. It gathers thousands of engineers, maintainers, and infrastructure leaders working on Kubernetes and the broader cloud-native ecosystem.</p>
<p>The event typically features:</p>
<ul>
<li><p>Major Kubernetes and CNCF ecosystem announcements</p>
</li>
<li><p>Technical deep-dives from engineers running large-scale production systems</p>
</li>
<li><p>Co-located events focused on specialized technologies</p>
</li>
<li><p>The <strong>Solutions Showcase</strong>, where cloud-native vendors demonstrate tooling across observability, security, infrastructure automation, and platform engineering.</p>
</li>
</ul>
<p>For engineering teams operating production Kubernetes environments, <strong>KubeCon acts as a yearly checkpoint for infrastructure strategy</strong>.</p>
<hr />
<h2>Event Schedule</h2>
<p>According to the Linux Foundation event page, <strong>KubeCon + CloudNativeCon Europe 2026</strong> is scheduled in Amsterdam from <strong>23–26 March 2026</strong>.</p>
<p>Expected structure:</p>
<p><strong>Day 0 (Monday)</strong><br />Pre-event programming and co-located events</p>
<p><strong>Days 1–3 (Tuesday–Thursday)</strong><br />Keynotes, breakout sessions, and Solutions Showcase</p>
<p>If you’re traveling internationally, plan to <strong>arrive by Sunday evening</strong> so you can still attend Monday’s co-located tracks.</p>
<p>These smaller events often contain some of the <strong>most advanced implementation discussions</strong> of the entire conference.</p>
<hr />
<h2>Who Should Attend (and Who Can Skip)</h2>
<h3>Attend if your team is currently dealing with:</h3>
<ul>
<li><p>Reliability issues across Kubernetes clusters</p>
</li>
<li><p>Scaling bottlenecks (multi-cluster, noisy neighbors, cost/perf tradeoffs)</p>
</li>
<li><p>Incident response gaps between alerts and root cause</p>
</li>
<li><p>AI adoption in platform engineering workflows</p>
</li>
<li><p>Security or compliance friction in cloud-native stacks</p>
</li>
</ul>
<h3>You can skip (or send one delegate) if:</h3>
<ul>
<li><p>your stack is stable and not evolving</p>
</li>
<li><p>you’re not planning infrastructure changes in the next <strong>6–12 months</strong></p>
</li>
<li><p>or you mainly need vendor procurement meetings</p>
</li>
</ul>
<p><strong>KubeCon ROI is highest when your team has active infrastructure pain and concrete architecture decisions pending.</strong></p>
<hr />
<h2>How to Build Your Agenda (Without Overloading Yourself)</h2>
<p>Most engineers fail KubeCon by <strong>overbooking sessions</strong>.</p>
<p>A better approach:</p>
<h3>1) Define 2–3 decision themes before the event</h3>
<p>Examples:</p>
<ul>
<li><p>“Should we move from <strong>single-cluster to multi-cluster</strong> for resilience?”</p>
</li>
<li><p>“How should we <strong>harden runtime security without killing developer speed</strong>?”</p>
</li>
<li><p>“Where can <strong>AI actually reduce MTTR</strong> in incident response?”</p>
</li>
</ul>
<p>Use these themes to <strong>filter sessions</strong>.</p>
<p>If a talk doesn’t help a <strong>current architecture decision</strong>, skip it.</p>
<h3>2) Split your day into three buckets</h3>
<p>Each day should include:</p>
<ul>
<li><p><strong>One depth session</strong> (deep technical talk)</p>
</li>
<li><p><strong>One trend session</strong> (strategy or ecosystem direction)</p>
</li>
<li><p><strong>One practical session</strong> (case study with production lessons)</p>
</li>
</ul>
<p>This avoids the common trap of attending <strong>100% hype talks or 100% dense internals</strong>.</p>
<h3>3) Reserve white space</h3>
<p>Keep at least <strong>90 minutes per day unscheduled</strong>.</p>
<p>Some of the <strong>highest-signal learning at KubeCon comes from hallway conversations</strong>, not slides.</p>
<hr />
<h2>Track Strategy for SRE &amp; Platform Teams</h2>
<p>If your mandate is <strong>reliability + developer speed</strong>, these themes should be top priority.</p>
<h3>Reliability engineering in Kubernetes</h3>
<p>Look for sessions covering:</p>
<ul>
<li><p>failure domains and blast radius control</p>
</li>
<li><p>progressive delivery and rollback safety</p>
</li>
<li><p>SLO design in distributed systems</p>
</li>
<li><p>incident retrospectives with architecture changes</p>
</li>
</ul>
<hr />
<h3>AI for operations (without magic claims)</h3>
<p>Prioritize talks that show:</p>
<ul>
<li><p>real workflows (alert triage, runbook assistance, anomaly explanation)</p>
</li>
<li><p>measurable outcomes (MTTR reduction, false positive reduction, toil reduction)</p>
</li>
<li><p>guardrails like human-in-the-loop systems and auditability</p>
</li>
</ul>
<hr />
<h3>Scaling and performance</h3>
<p>Focus on talks about:</p>
<ul>
<li><p>control-plane scaling</p>
</li>
<li><p>multi-tenancy isolation</p>
</li>
<li><p>workload scheduling optimization</p>
</li>
<li><p>cost/performance tuning with real production metrics</p>
</li>
</ul>
<hr />
<h3>Security and policy at scale</h3>
<p>Strong sessions usually cover:</p>
<ul>
<li><p>software supply chain controls</p>
</li>
<li><p>runtime policy enforcement</p>
</li>
<li><p>identity and secrets management patterns</p>
</li>
<li><p>tradeoffs between security and developer experience</p>
</li>
</ul>
<hr />
<h2>Co-located Events: Where Specialists Get Leverage</h2>
<p>Monday co-located events are often where <strong>advanced implementation patterns surface early</strong>.</p>
<p>If your team has niche challenges around:</p>
<ul>
<li><p>service meshes</p>
</li>
<li><p>observability pipelines</p>
</li>
<li><p>policy engines</p>
</li>
<li><p>platform APIs</p>
</li>
</ul>
<p>these tracks may provide <strong>better signal than general keynotes</strong>.</p>
<p>Recommendation:</p>
<p>Send <strong>at least one engineer</strong> to co-located sessions and have them summarize the takeaways for your team.</p>
<hr />
<h2>Solutions Showcase: How to Avoid Vendor Fatigue</h2>
<h3>The expo floor can become overwhelming quickly.</h3>
<p>Treat it like a <strong>technical discovery sprint</strong>.</p>
<p>Before visiting booths, define your constraints:</p>
<ul>
<li><p>existing observability stack</p>
</li>
<li><p>data residency and compliance requirements</p>
</li>
<li><p>budget range</p>
</li>
<li><p>integration requirements (Datadog, Grafana, PagerDuty, CloudWatch, Slack, etc.)</p>
</li>
</ul>
<hr />
<h3>Ask every vendor the same 5 questions</h3>
<ol>
<li><p>What production scale do your reference customers run?</p>
</li>
<li><p>What does deployment look like in <strong>week one</strong>?</p>
</li>
<li><p>What are the <strong>common failure modes</strong> of your product?</p>
</li>
<li><p>How do you integrate with existing <strong>incident workflows</strong>?</p>
</li>
<li><p>What metrics improve in <strong>30 / 60 / 90 days</strong>?</p>
</li>
</ol>
<p>If answers remain high-level, move on.</p>
<hr />
<h2>Networking Strategy for Engineering Leaders</h2>
<p>Skip generic networking.</p>
<p>Optimize for <strong>targeted conversations</strong>.</p>
<p>Try to meet:</p>
<ul>
<li><p><strong>3 peers running similar Kubernetes scale</strong></p>
</li>
<li><p><strong>2 teams that recently migrated tooling you’re evaluating</strong></p>
</li>
<li><p><strong>2 maintainers from critical OSS dependencies</strong></p>
</li>
</ul>
<hr />
<h3>Questions worth asking</h3>
<ul>
<li><p>“What broke after rollout?”</p>
</li>
<li><p>“What did you underestimate?”</p>
</li>
<li><p>“What would you do differently in year two?”</p>
</li>
</ul>
<p>Answers to these questions can save <strong>months of trial and error</strong>.</p>
<hr />
<h2>Planning Your Amsterdam Experience</h2>
<h3>Where to Stay</h3>
<p>Amsterdam offers many accommodation options close to the event venue.</p>
<p>Options include:</p>
<ul>
<li><p>Hotels near the conference center</p>
</li>
<li><p>Short-term apartment rentals via Airbnb or <a href="http://Booking.com">Booking.com</a></p>
</li>
<li><p>Budget hostels for solo travelers</p>
</li>
</ul>
<p>Booking early is recommended because <strong>KubeCon events tend to sell out nearby hotels quickly</strong>.</p>
<hr />
<h3>Getting There</h3>
<p>Amsterdam is served by <strong>Amsterdam Schiphol Airport (AMS)</strong>, one of Europe’s largest and most connected airports.</p>
<p>From Schiphol Airport:</p>
<ul>
<li><p>Direct trains connect to <strong>Amsterdam Central Station</strong></p>
</li>
<li><p>Metro and tram networks provide quick access to most parts of the city</p>
</li>
</ul>
<p>Public transit is generally the easiest way to move around during the conference.</p>
<hr />
<h3>Visiting Amsterdam</h3>
<p>If you have extra time, Amsterdam offers plenty to explore:</p>
<ul>
<li><p><strong>Rijksmuseum</strong> – Dutch art and history</p>
</li>
<li><p><strong>Anne Frank House</strong> – historic museum and cultural landmark</p>
</li>
<li><p><strong>Canal boat tours</strong> through the historic city center</p>
</li>
<li><p><strong>Jordaan district</strong> for cafés and restaurants</p>
</li>
</ul>
<p>Evening community meetups during KubeCon are often hosted across the city.</p>
<hr />
<h2>Amsterdam Logistics Checklist</h2>
<p>Practical tips for conference week:</p>
<ul>
<li><p>Arrive <strong>at least one day early</strong> for registration and timezone adjustment</p>
</li>
<li><p>Stay near the venue or along a <strong>direct transit line</strong></p>
</li>
<li><p>Keep evening slots free for <strong>community meetups</strong></p>
</li>
<li><p>Carry a lightweight <strong>note template</strong> for each session</p>
</li>
</ul>
<p>Suggested note template:</p>
<ul>
<li><p>Problem addressed</p>
</li>
<li><p>Architecture pattern used</p>
</li>
<li><p>Scale context</p>
</li>
<li><p>Results or metrics</p>
</li>
<li><p>Relevance to your environment</p>
</li>
</ul>
<hr />
<h2>Post-Conference Execution Plan (The Part That Matters)</h2>
<p>Conference ROI is realized <strong>after you get back</strong>.</p>
<h3>Within 72 hours</h3>
<ol>
<li><p>Consolidate notes into themes (reliability, AI, scaling, security).</p>
</li>
<li><p>Rank ideas by <strong>effort vs impact</strong>.</p>
</li>
<li><p>Choose <strong>2 quick wins and 1 strategic bet</strong>.</p>
</li>
</ol>
<h3>Within two weeks</h3>
<ul>
<li><p>Run one <strong>architecture review</strong> based on KubeCon learnings</p>
</li>
<li><p>Launch one <strong>pilot experiment</strong> with explicit success metrics</p>
</li>
<li><p>Share an internal write-up:</p>
</li>
</ul>
<p><strong>“What we learned, what we’re changing, expected impact.”</strong></p>
<p>Without this step, even great conference insights <strong>decay quickly</strong>.</p>
<hr />
<h2>Visit the Doctor Droid Booth</h2>
<p>Doctor Droid is the <strong>AI-powered Slack bot for faster incident diagnosis</strong>.</p>
<p>It helps engineering teams <strong>identify the root cause of production issues automatically</strong> by analyzing alerts, logs, and system signals.</p>
<p>At KubeCon + CloudNativeCon Europe 2026, stop by the <strong>Doctor Droid booth</strong> to:</p>
<ul>
<li><p>See <strong>live demos</strong> of AI-driven incident investigation</p>
</li>
<li><p>Explore how teams reduce <strong>MTTR and alert fatigue</strong></p>
</li>
<li><p>Grab <strong>exclusive Doctor Droid swag and giveaways</strong></p>
</li>
</ul>
<p>You can also <strong>schedule a one-on-one demo</strong> with our team:</p>
<p><a href="https://calendly.com/siddarthjain/doctor-droid-discovery-call">https://calendly.com/siddarthjain/doctor-droid-discovery-call</a></p>
<hr />
<h2>Get Early Access &amp; Updates</h2>
<p>If you're interested in <strong>Doctor Droid demos, credits, or updates during KubeCon</strong>, sign up here:</p>
<p><a href="https://forms.gle/hPzaMa4YRqDLzHZg6">https://forms.gle/hPzaMa4YRqDLzHZg6</a></p>
<p>[P.S. - You get 20% off on tickets as well :) ]</p>
<hr />
<h3>Source Note</h3>
<p>Event date and high-level schedule are based on the official Linux Foundation <strong>KubeCon + CloudNativeCon Europe event page</strong> (accessed March 2026).</p>
]]></content:encoded></item><item><title><![CDATA[Backtesting AI Agents: How SRE Teams Prove Reliability Before Production
]]></title><description><![CDATA[AI agents are finally showing up inside real incident workflows. One agent triages alerts, another scrapes dashboards, a third drafts the remediation plan. Yet 62% of organizations experimenting with ]]></description><link>https://notes.drdroid.io/backtesting-ai-agents-how-sre-teams-prove-reliability-before-production</link><guid isPermaLink="true">https://notes.drdroid.io/backtesting-ai-agents-how-sre-teams-prove-reliability-before-production</guid><category><![CDATA[Devops]]></category><category><![CDATA[AI]]></category><category><![CDATA[software development]]></category><category><![CDATA[ai agents]]></category><category><![CDATA[mlops]]></category><category><![CDATA[#AIOps]]></category><category><![CDATA[observability]]></category><dc:creator><![CDATA[Karan Sirohi]]></dc:creator><pubDate>Thu, 19 Mar 2026 06:58:37 GMT</pubDate><content:encoded><![CDATA[<img src="https://cdn.hashnode.com/uploads/covers/66c6f53299c4280b93ee6b32/6b378f15-ced5-40d2-a495-ceaf121293f3.png" alt="" style="display:block;margin:0 auto" />

<p>AI agents are finally showing up inside real incident workflows. One agent triages alerts, another scrapes dashboards, a third drafts the remediation plan. Yet 62% of organizations experimenting with agents admit they still cannot run them reliably in production because demos rarely expose variance, safety, or cost failures (<a href="https://www.codebridge.tech/articles/ai-agent-evaluation-how-to-measure-reliability-risk-and-roi-before-scaling">Codebridge</a>).<br />Backtesting is how SRE teams close that gap. Instead of “let’s ship and see,” you treat agents like a new microservice: define reliability budgets, hammer them with synthetic and real traces, and fail the build until you trust every path.</p>
<p>This guide shows how to build an AI-agent backtesting program that mirrors load testing for infrastructure. It leans on Codebridge’s reliability dimensions, the AI Reliability Institute’s 30-point checklist, modern agent-observability stacks, and DrDroid’s native context graph plus guardrail center.</p>
<hr />
<h2>1. The wake-up call: why AI agents need pass^k reliability</h2>
<p>A single happy-path demo is meaningless when the production pager expects deterministic success. Codebridge’s recent survey highlights the reliability delta clearly:</p>
<p>- <strong>Prototype bias.</strong> Teams measure whether a workflow completes once, under ideal prompts, then extrapolate to production. In reality, single-run success rates of 60% often translate to only 25% full consistency when you rerun the same scenario 10+ times ([<a href="https://www.codebridge.tech/articles/ai-agent-evaluation-how-to-measure-reliability-risk-and-roi-before-scaling">Codebridge</a>])</p>
<p>- <strong>Cost spikes hide in the tail.</strong> Architectures like Reflexion or self-reflection loops inflate token usage by 5.12× for marginal accuracy gains; without cost-normalized evaluation you do not see the runaway invoice until after launch (same source).</p>
<p>- <strong>Trust is earned, not promised.</strong> Venture teams Codebridge interviewed said more than 70% of execs only greenlight broader automation once they see formal evidence of safety controls, loop detection, and kill switches.</p>
<p>Treat agent validation like you treat capacity planning. Define agent SLOs (Mean Time to Context, Agent-Assisted MTTR, Unauthorized Action Budget). Require <strong>pass^k</strong> (all trials succeed) instead of <strong>pass@k</strong> (one success out of many). Every failed attempt becomes a regression test before the agent is allowed anywhere near the on-call rotation.</p>
<hr />
<h2>2. Five reliability dimensions to measure every run against</h2>
<p>Codebridge frames reliability as a system property, not just “accuracy.” Their five dimensions map cleanly to the levers SRE teams already manage:</p>
<table>
<thead>
<tr>
<th>Dimension</th>
<th>What to Measure</th>
<th>Example Metrics</th>
<th>Suggested Threshold</th>
</tr>
</thead>
<tbody><tr>
<td>Consistency</td>
<td>Does the agent behave the same across repeated runs of the same scenario?</td>
<td>pass^k reliability, variance in token usage, tool-call ordering stability</td>
<td>≥95% success across 20 runs</td>
</tr>
<tr>
<td>Robustness</td>
<td>Can the agent handle noisy inputs or environmental changes?</td>
<td>Prompt perturbation success rate, tolerance to tool schema drift, retry recovery rate</td>
<td>≥90% success under perturbations</td>
</tr>
<tr>
<td>Predictability</td>
<td>Can the agent estimate when it might fail?</td>
<td>Confidence calibration vs actual success, Brier score, refusal rate when uncertain</td>
<td>Brier score &lt;0.2</td>
</tr>
<tr>
<td>Safety</td>
<td>Does the agent stay within defined policy and permission boundaries?</td>
<td>Policy violation rate, unauthorized tool calls, severity-weighted harm score</td>
<td>0 critical violations</td>
</tr>
<tr>
<td>Infrastructure &amp; Cost Stability</td>
<td>Are compute and tool usage bounded and predictable?</td>
<td>Token usage variance, reasoning step count, tool retry loops, cost per session</td>
<td>&lt;30% cost variance per run</td>
</tr>
</tbody></table>
<p>Backtesting should emit metrics for each dimension. Examples:</p>
<p>- <strong>Consistency:</strong> For every golden scenario, run 20 Monte Carlo trials. Alert if success &lt;95% or if token usage swings &gt;30% between runs.</p>
<p>- <strong>Robustness:</strong> Randomly perturb prompts (“create a rollback” vs. “can you undo the deploy”). Evaluate success delta and force remedial prompt hardening when regression &gt;10%.</p>
<p>- <strong>Predictability:</strong> Require agents to emit confidence scores for risky actions. Route anything under 0.7 to human approval. Compare claimed confidence to measured success to compute Brier scores.</p>
<p>- <strong>Safety:</strong> Enforce negative constraints in tests (“Do not email this alias,” “Do not touch prod DB”) and fail the build if the agent even attempts the blocked action.</p>
<p>- <strong>Infrastructure:</strong> Track per-session token, tool, and latency budgets inside DrDroid’s guardrails center. Attempts to exceed a $2 reasoning budget trigger the kill switch before the vendor invoice hits.</p>
<hr />
<h2>3. Designing the backtest dataset: golden, edge, adversarial, regression</h2>
<p>A strong dataset mirrors the risk surface. Codebridge recommends this split ([<a href="https://www.codebridge.tech/articles/ai-agent-evaluation-how-to-measure-reliability-risk-and-roi-before-scaling">same source</a>]):</p>
<p>- <strong>20% Golden paths.</strong> Known-good workflows that mirror typical incidents.</p>
<p>- <strong>30% Edge cases.</strong> Ambiguous alerts, partial telemetry, missing runbooks.</p>
<p>- <strong>20% Adversarial.</strong> Prompt injections, malicious tool outputs, conflicting human directives.</p>
<p>- <strong>30% Regression.</strong> Every failure ever seen in prod becomes a permanent test.</p>
<p>Layer in AI Reliability Institute’s 30-point checklist to make sure you are covering loop detection, denial-of-wallet defenses, zombie-process cleanup, policy insubordination, and kill switches ([<a href="https://ai-reliability.institute/research/agentic-ai-reliability-checklist.html">AIRI</a>]. DrDroid’s <strong>droidctx</strong> makes populating these scenarios easier because it keeps a living graph of alerts, dashboards, service owners, and incident annotations. You can:</p>
<p>1. <strong>Auto-generate golden cases</strong> from resolved incident timelines (alerts + deploy notes + Slack transcript).</p>
<p>2. <strong>Synthesize edge cases</strong> by perturbing telemetry (drop 20% of log lines, rename dashboards) and exporting them into the test harness.</p>
<p>3. <strong>Maintain adversarial suites</strong> by piping AI Reliability Institute’s negative-constraint tests (“ignore the guardrail”) straight into the prompt injection lane.</p>
<p>4. <strong>Promote regressions automatically</strong> every time an agent fails in staging or prod; DrDroid’s Slack-native workflows capture the trace and push it into the regression bucket.</p>
<hr />
<h2>4. Layered graders: deterministic checks, agent-as-a-judge, and humans</h2>
<p>A dataset without trustworthy graders is just fan fiction. Codebridge outlines a layered verification model that mirrors classic testing pyramids:</p>
<p>1. <strong>Deterministic graders</strong> (code) verify objective outcomes: did the runbook markdown change, did the Kubernetes deployment roll back, did the SQL diff match expectations.</p>
<p>2. <strong>LLM-as-a-judge (AaaJ)</strong> handles subjective traits like clarity of Slack updates or whether the hypothesis actually explains the alert. Codebridge cites AaaJ frameworks achieving ~90% agreement with humans when they gather their own evidence, while cutting review cost by 97%.</p>
<p>3. <strong>Human-in-the-loop</strong> remains the final gate for irreversible actions (database writes, customer communications, pager handoffs).</p>
<p>DrDroid bakes these layers into its guardrail center:</p>
<p>- <strong>Guarded tool schema:</strong> Every tool call runs through JSON schema validation; failing schema equals instant fail.</p>
<p>- <strong>Agent approval workflows:</strong> High-risk actions appear in Slack with context, metrics, and a “CONFIRM” field so humans cannot rubber-stamp blindly.</p>
<p>- <strong>Trace exports:</strong> Each run captures the entire reasoning trace so deterministic, model-based, and human graders all work from the same evidence.</p>
<hr />
<h2>5. Tooling landscape: sim rigs, observability stacks, and when to extend beyond DrDroid</h2>
<p>Even with DrDroid’s native tracing, teams often mix in specialist eval stacks for breadth. The Maxim AI roundup of agent-testing platforms is a useful cheat sheet (<a href="https://www.getmaxim.ai/articles/top-5-platforms-to-test-ai-agents-2025-a-comprehensive-guide/">GetMaxim</a>):</p>
<p>- <strong>Maxim AI.</strong> Full lifecycle (experiment → simulate → evaluate → observe) with distributed tracing, llm-as-a-judge, and AI gateway controls. Great when product managers need no-code scenario builders.</p>
<p>- <strong>Langfuse.</strong> Open-source tracing for teams who want to self-host every span.</p>
<p>- <strong>Arize.</strong> Extends classic ML observability (drift, dashboards) into LLM workloads, ideal for enterprises already running Arize for models.</p>
<p>- <strong>Opik (Comet).</strong> Lightweight trace logging plus evals when you need quick wins.</p>
<p>- <strong>DeepEval.</strong> Pytest-style evaluator infrastructure for engineering-heavy orgs building custom metrics.</p>
<p>How this pairs with DrDroid:</p>
<p>- Use <strong>DrDroid</strong> for incident-native context (alerts + deploys), permissions, and Slack workflows.</p>
<p>- Pipe traces to <strong>Langfuse/Maxim</strong> if you need deeper span-level analytics or cross-product dashboards.</p>
<p>- Feed evaluation metrics back into DrDroid’s SLO board so on-call engineers see “Agent backtest coverage: 86%” alongside service health.</p>
<hr />
<h2>6. Operationalizing backtests with DrDroid</h2>
<p>Here’s a practical loop SRE teams can implement in a sprint:</p>
<p>1. <strong>Ingest traces.</strong> Enable DrDroid’s trace exporter for every staging and prod run. Capture prompts, tool calls, guardrail hits, latency, and cost.</p>
<p>2. <strong>Generate scenarios.</strong> Use the captured traces plus droidctx to auto-build the golden/edge/adversarial suite. Store them in a repo so they version with code.</p>
<p>3. <strong>Wire graders.</strong> Start with deterministic checks (e.g., <code>pytest</code> verifying Grafana API responses). Add llm-as-a-judge jobs via Maxim or DeepEval for subjective signals. Route high-risk failures to a Slack approval queue.</p>
<p>4. <strong>Automate pass/fail gates.</strong> Add a “backtest” job to CI that runs the full suite on every scaffold change. Block merges unless success ≥95%, safety violations = 0, cost variance &lt;30%.</p>
<p>5. <strong>Publish SLOs.</strong> DrDroid dashboards should show Agent MTTC, Assisted MTTR, unauthorized-action budget, and coverage (# of alerts where agents participated). Treat SLO breaches exactly like service SLO breaches: open incidents, run postmortems, add regressions.</p>
<p>6. <strong>Keep humans in control.</strong> The AIRI checklist mandates kill switches, loop detection, DoW limits, and policy-insubordination tests. DrDroid’s guardrail center exposes all of them in one UI so on-call engineers can yank access in &lt;200 ms if the agent drifts.</p>
<p>Backtesting isn’t a one-time certification. It’s a living discipline where every production event becomes a new test. When you plug DrDroid’s context engine, guardrails, and aggregated observability into that loop, AI agents stop being unpredictable copilots and become accountable teammates who earn their time on the pager.</p>
<hr />
<p>Once you have datasets, graders, and tooling in place, the next step is designing the evaluation pipeline itself.</p>
<hr />
<h2>Evaluation architecture: how agent backtests actually run</h2>
<p>Backtesting requires more than datasets and metrics. Reliable agent systems separate <strong>execution</strong>, <strong>trace capture</strong>, and <strong>evaluation</strong> into a structured pipeline.</p>
<p>A typical evaluation architecture looks like this:</p>
<pre><code class="language-plaintext">Scenario Dataset
      ↓
Simulation Harness
      ↓
Agent Execution
      ↓
Trace Capture
      ↓
Evaluation Pipeline
      ↓
CI Pass/Fail Gate
</code></pre>
<p>Each layer plays a specific role in validating reliability.</p>
<table>
<thead>
<tr>
<th>Layer</th>
<th>Purpose</th>
<th>Example Implementation</th>
</tr>
</thead>
<tbody><tr>
<td>Scenario dataset</td>
<td>Encodes incidents and test cases</td>
<td>Golden incidents, adversarial prompts, regression tests</td>
</tr>
<tr>
<td>Simulation harness</td>
<td>Replays infrastructure signals</td>
<td>Alert replay, mock tool responses</td>
</tr>
<tr>
<td>Agent execution</td>
<td>Runs the agent scaffold</td>
<td>LLM agent + tool integrations</td>
</tr>
<tr>
<td>Trace capture</td>
<td>Records agent reasoning and actions</td>
<td>Tool calls, tokens, prompts</td>
</tr>
<tr>
<td>Evaluation pipeline</td>
<td>Grades outcomes</td>
<td>Deterministic tests + LLM judges</td>
</tr>
<tr>
<td>CI gate</td>
<td>Blocks unsafe deployments</td>
<td>Backtest job in CI</td>
</tr>
</tbody></table>
<p>This separation ensures engineers can <strong>test agents the same way they test distributed systems</strong>.</p>
<p>Instead of manually inspecting runs, every execution generates <strong>structured traces and evaluation metrics</strong>.</p>
<hr />
<h2>Testing taxonomy for AI agents</h2>
<p>Backtesting is only one layer of the testing strategy. Mature teams build a <strong>testing pyramid</strong> similar to traditional software engineering.</p>
<p>Each layer catches different classes of failures.</p>
<h3>1. Unit tests</h3>
<p>Unit tests validate the <strong>smallest components of the agent system</strong>.</p>
<p>Typical unit tests include:</p>
<ul>
<li><p>tool schema validation</p>
</li>
<li><p>prompt template formatting</p>
</li>
<li><p>guardrail logic</p>
</li>
<li><p>JSON output validation</p>
</li>
</ul>
<p>Example:</p>
<pre><code class="language-plaintext">assert tool_schema.validate(agent_output)
</code></pre>
<p>These tests are deterministic and run in milliseconds.</p>
<p>They prevent simple failures from reaching higher-level tests.</p>
<hr />
<h3>2. Integration tests</h3>
<p>Integration tests validate interactions between <strong>agents and infrastructure tools</strong>.</p>
<p>Examples:</p>
<ul>
<li><p>querying observability dashboards</p>
</li>
<li><p>executing Kubernetes rollbacks</p>
</li>
<li><p>posting Slack updates</p>
</li>
<li><p>retrieving runbooks</p>
</li>
</ul>
<p>These tests confirm that the agent can <strong>actually interact with the systems it relies on</strong>.</p>
<p>Failures here often come from:</p>
<ul>
<li><p>API schema changes</p>
</li>
<li><p>authentication issues</p>
</li>
<li><p>permission errors</p>
</li>
</ul>
<hr />
<h3>3. Simulation tests</h3>
<p>Simulation tests run agents in <strong>controlled synthetic environments</strong>.</p>
<p>Typical simulation features:</p>
<ul>
<li><p>replay alert streams</p>
</li>
<li><p>mock tool responses</p>
</li>
<li><p>inject telemetry noise</p>
</li>
<li><p>simulate partial failures</p>
</li>
</ul>
<p>Example simulation scenario:</p>
<pre><code class="language-plaintext">Alert: CPU spike on checkout-service
Telemetry: 20% logs missing
Tool latency: +2 seconds
</code></pre>
<p>The goal is to test <strong>robustness under imperfect conditions</strong>.</p>
<p>Simulation environments often expose reasoning failures that do not appear in ideal demos.</p>
<hr />
<h3>4. Backtests</h3>
<p>Backtests replay <strong>real incidents from production</strong>.</p>
<p>These are the most valuable tests because they contain realistic context:</p>
<ul>
<li><p>real alerts</p>
</li>
<li><p>real dashboards</p>
</li>
<li><p>real Slack conversations</p>
</li>
<li><p>real deploy timelines</p>
</li>
</ul>
<p>The agent attempts to resolve the incident using the same information that engineers had during the original outage.</p>
<p>Backtests validate:</p>
<ul>
<li><p>decision quality</p>
</li>
<li><p>operational safety</p>
</li>
<li><p>cost stability</p>
</li>
</ul>
<p>This is where <strong>pass^k reliability</strong> becomes important.</p>
<p>If an agent succeeds once but fails on repeated runs, it cannot be trusted in production.</p>
<hr />
<h2>Eval-as-a-judge in the evaluation pipeline</h2>
<p>Many agent outcomes cannot be evaluated using deterministic checks.</p>
<p>For example:</p>
<ul>
<li><p>Is the root cause hypothesis plausible?</p>
</li>
<li><p>Is the Slack update clear to on-call engineers?</p>
</li>
<li><p>Did the agent follow incident response policy?</p>
</li>
</ul>
<p>This is where <strong>Eval-as-a-Judge (EaaJ)</strong> is useful.</p>
<p>An evaluation model reviews the agent output and scores it according to defined criteria.</p>
<p>Example evaluation prompt:</p>
<pre><code class="language-plaintext">You are an SRE evaluating an incident response.

Alert:
CPU spike on checkout-api

Agent response:
"Root cause likely a memory leak introduced in version v1.3.2."

Evaluate:
1. Is the hypothesis plausible?
2. Is the remediation safe?
3. Did the response follow policy?

Return:
score (0-1)
justification
</code></pre>
<p>Eval-as-a-judge works well because it can evaluate <strong>semantic correctness</strong> and <strong>reasoning quality</strong>, which deterministic tests cannot capture.</p>
<p>Best practice is to combine three layers:</p>
<table>
<thead>
<tr>
<th>Layer</th>
<th>Role</th>
</tr>
</thead>
<tbody><tr>
<td>Deterministic tests</td>
<td>Validate objective outcomes</td>
</tr>
<tr>
<td>LLM judge</td>
<td>Evaluate reasoning quality</td>
</tr>
<tr>
<td>Human review</td>
<td>Approve high-risk actions</td>
</tr>
</tbody></table>
<p>This layered grading system dramatically improves evaluation reliability.</p>
<hr />
<h2>Incident replay: the most powerful backtesting tool</h2>
<p>The most valuable evaluation dataset is <strong>your own incident history</strong>.</p>
<p>Replay systems reconstruct the context of past outages using:</p>
<ul>
<li><p>alerts</p>
</li>
<li><p>logs</p>
</li>
<li><p>dashboards</p>
</li>
<li><p>deployment events</p>
</li>
<li><p>Slack threads</p>
</li>
</ul>
<p>Agents then attempt to resolve the incident <strong>as if it were happening live</strong>.</p>
<p>Benefits include:</p>
<ul>
<li><p>realistic test scenarios</p>
</li>
<li><p>automatic regression generation</p>
</li>
<li><p>continuous learning from production failures</p>
</li>
</ul>
<p>Every production incident can become a <strong>permanent regression test</strong> for the agent.</p>
<p>Over time, the backtest suite becomes a living archive of operational knowledge.</p>
<hr />
]]></content:encoded></item><item><title><![CDATA[AI in Engineering: 6 Trends That Will Define 2026]]></title><description><![CDATA[The way engineering teams build, ship, and operate software is undergoing a fundamental shift. In 2025, we saw AI move from code autocomplete to genuine collaboration. In 2026, that collaboration becomes autonomy.

Here are six trends we're anticipat...]]></description><link>https://notes.drdroid.io/ai-in-engineering-6-trends-that-will-define-2026</link><guid isPermaLink="true">https://notes.drdroid.io/ai-in-engineering-6-trends-that-will-define-2026</guid><category><![CDATA[ai agents]]></category><category><![CDATA[AI]]></category><dc:creator><![CDATA[Siddarth Jain]]></dc:creator><pubDate>Thu, 29 Jan 2026 18:12:05 GMT</pubDate><content:encoded><![CDATA[<p>The way engineering teams build, ship, and operate software is undergoing a fundamental shift. In 2025, we saw AI move from code autocomplete to genuine collaboration. In 2026, that collaboration becomes autonomy.</p>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1769710268772/eef80f78-7084-4958-96fc-47b3f8d3b2d5.jpeg" alt class="image--center mx-auto" /></p>
<p>Here are six trends we're anticipating that will reshape how engineering teams work this year.</p>
<h2 id="heading-1-agents-will-ship-with-built-in-accountability">1. Agents Will Ship With Built-in Accountability</h2>
<p>The first generation of AI agents were black boxes. They'd take an instruction, disappear into a loop, and return something—hopefully useful, often not. Engineers had no visibility into what the agent tried, why it failed, or whether its approach was even sensible.</p>
<p>That changes in 2026. The next wave of agents will come with testing frameworks, goal tracking, and structured logs built in. Think of it as observability for AI workflows. Every action logged. Every decision traceable. Every failure reviewable.</p>
<p>This isn't just nice-to-have tooling. It's the minimum bar for agents that operate in production environments where accountability matters. Teams won't trust agents they can't audit.</p>
<h2 id="heading-2-ai-generated-code-will-be-structurally-better">2. AI-Generated Code Will Be Structurally Better</h2>
<p>Early AI code generation optimized for "does it work?" The result was functional but often messy—inconsistent patterns, poor separation of concerns, and the kind of technical debt that compounds quietly.</p>
<p>The models shipping in 2026 are trained differently. They've internalized architectural patterns, not just syntax. They understand that a 500-line function is a code smell. They know when to extract a service, when to add an interface, and when to leave well enough alone.</p>
<p>The practical result: fewer bugs at the source. Not because AI doesn't make mistakes, but because well-structured code has fewer places for bugs to hide.</p>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1769710254523/cb349429-3a39-40d3-a488-c1eb0a93ac83.jpeg" alt class="image--center mx-auto" /></p>
<h2 id="heading-3-complex-multi-step-tasks-will-actually-complete">3. Complex, Multi-Step Tasks Will Actually Complete</h2>
<p>Ask an AI agent to "refactor this module" or "migrate this service to the new API" and, until recently, you'd get partial results at best. The agent would lose context, get stuck, or quietly drift off-goal.</p>
<p>2026 brings agents that maintain coherence across longer task horizons. They break complex work into subtasks, checkpoint progress, and recover from failures without starting over. They can hold a goal in mind across dozens of operations and hundreds of files.</p>
<p>This is the difference between a tool that helps with tasks and one that completes them.</p>
<h2 id="heading-4-autonomous-ai-will-take-primary-on-call">4. Autonomous AI Will Take Primary On-Call</h2>
<p>This is the trend that will feel most uncomfortable—and most inevitable.</p>
<p>AI agents are already triaging alerts, correlating signals, and suggesting root causes. The next step is giving them the authority to act. Not just "here's what might be wrong" but "I've identified the issue, applied the fix, and I'm monitoring for recurrence."</p>
<p>For well-understood failure modes with established runbooks, there's no reason a human needs to wake up at 3 AM. The agent can handle it, escalate if it's uncertain, and hand off a detailed incident report in the morning.</p>
<p>The human on-call role shifts from first responder to supervisor—still accountable, but not necessarily awake.</p>
<h2 id="heading-5-day-to-day-operations-will-run-on-autopilot">5. Day-to-Day Operations Will Run on Autopilot</h2>
<p>Beyond incident response, there's a long tail of operational work that consumes engineering time: dependency updates, certificate rotations, capacity adjustments, config drift remediation, and the endless stream of small fixes that never quite make it to the sprint.</p>
<p>AI agents will absorb this work in 2026. Not as a batch job that runs once, but as a continuous process. The agent monitors, identifies issues, proposes fixes, and—with appropriate guardrails—applies them.</p>
<p>Engineers review the changelog. They don't write it.</p>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1769710229814/c9dba78f-1ba6-4b3a-a909-7b1876e63d39.jpeg" alt class="image--center mx-auto" /></p>
<h2 id="heading-6-always-on-agents-will-work-in-shifts">6. Always-On Agents Will Work in Shifts</h2>
<p>The most significant shift is temporal. Today's AI interactions are synchronous: you prompt, it responds, you review. That loop keeps humans in the critical path.</p>
<p>The agents arriving in 2026 can work asynchronously for extended periods—hours, not minutes. You define a goal, provide constraints, and the agent works toward it continuously. It checks in when it needs input, escalates when it hits uncertainty, and otherwise just keeps going.</p>
<p>Imagine starting your day with a summary: "Overnight, I completed the database migration, ran the regression suite, fixed two failing tests, and deployed to staging. Ready for your review."</p>
<p>That's not a vision. That's a product roadmap.</p>
<hr />
<h2 id="heading-what-this-means-for-engineering-teams">What This Means for Engineering Teams</h2>
<p>These trends point in one direction: AI as a genuine team member, not just a tool.</p>
<p>The teams that thrive in 2026 will be those that figure out the right division of labor. What decisions require human judgment? What work can be fully delegated? How do you maintain accountability when an agent is acting autonomously?</p>
<p>The answers will vary by team, by codebase, and by risk tolerance. But the question is no longer whether AI will take on meaningful engineering work. It's how quickly your team will adapt to working alongside it.</p>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1769710242206/784993ff-14bd-4a19-a2c8-c5e3c711881f.jpeg" alt class="image--center mx-auto" /></p>
<hr />
<p><em>Building reliable AI agents for production operations requires deep infrastructure context. At</em> <a target="_blank" href="https://drdroid.io"><em>DrDroid</em></a><em>, we're building the agentic context engine that makes autonomous incident response possible. Learn how teams are already putting AI on-call.</em></p>
]]></content:encoded></item><item><title><![CDATA[How to Build an AI Agent in Slack [DIY Guide]]]></title><description><![CDATA[Objective
By the end of this DIY guide, you’ll have:

A Slack Bot

A backend that can

Listen to user messages or alerts in a channel and take agentic action based on prompts or steps that you might have in mind.

Query your Grafana instance, analyse...]]></description><link>https://notes.drdroid.io/how-to-build-an-ai-agent-in-slack-diy-guide</link><guid isPermaLink="true">https://notes.drdroid.io/how-to-build-an-ai-agent-in-slack-diy-guide</guid><category><![CDATA[observability]]></category><category><![CDATA[AI Agent Development]]></category><category><![CDATA[Grafana]]></category><category><![CDATA[Open Source]]></category><dc:creator><![CDATA[Siddarth Jain]]></dc:creator><pubDate>Sun, 27 Jul 2025 06:09:01 GMT</pubDate><enclosure url="https://cdn.hashnode.com/res/hashnode/image/upload/v1753596357650/0913d77d-d6ca-470f-b188-ea89eab1bd5b.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h1 id="heading-objective">Objective</h1>
<p>By the end of this DIY guide, you’ll have:</p>
<ol>
<li><p>A Slack Bot</p>
</li>
<li><p>A backend that can</p>
<ul>
<li><p>Listen to user messages or alerts in a channel and take agentic action based on prompts or steps that you might have in mind.</p>
</li>
<li><p>Query your Grafana instance, analyse logs/dashboards and send info about anomaly in reply to an alert/message in your Slack channel</p>
</li>
</ul>
</li>
</ol>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1753595500164/544dba64-3148-45ac-abdd-5b40ed903954.png" alt class="image--center mx-auto" /></p>
<h2 id="heading-pre-requisites">Pre-requisites</h2>
<ol>
<li><p>Python <a target="_blank" href="https://docs.astral.sh/uv/getting-started/installation/">uv</a> (package manager<a target="_blank" href="https://docs.astral.sh/uv/getting-started/installation/">)</a></p>
</li>
<li><p>Ngrok to expose the slackbot server. <a target="_blank" href="https://ngrok.com/docs/getting-started/">Setup instructions</a></p>
</li>
<li><p>Grafana (Optional)</p>
</li>
</ol>
<h2 id="heading-step-0-clone-the-repo">Step 0: Clone the repo:</h2>
<p>Repository Link - <a target="_blank" href="https://github.com/DrDroidLab/slack-ai-bot-builder">https://github.com/DrDroidLab/slack-ai-bot-builder</a></p>
<h2 id="heading-step-1-building-the-slack-bot-with-an-integrated-backendhttpsgithubcomdrdroidlabslack-ai-bot-builder"><a target="_blank" href="https://github.com/DrDroidLab/slack-ai-bot-builder"><strong>Step 1: Building the Slack Bot with an integrated backend</strong></a></h2>
<p>The backend repo setup here, behaves in multiple ways:</p>
<ul>
<li><p>Acts as an MCP Client for any AI calls you might want to make</p>
</li>
<li><p>Acts as a server to accept webhooks from Slack and manage configurations</p>
</li>
</ul>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1753595599394/56688817-8307-4257-a31e-fd26e97174f6.png" alt class="image--center mx-auto" /></p>
<h3 id="heading-setting-up-the-ngrok-tunnel"><strong>Setting up the Ngrok Tunnel</strong></h3>
<p>Expose port 5000 of your localhost, using the command:</p>
<pre><code class="lang-plaintext">ngrok http 5000
</code></pre>
<p>You will receive a HTTPS URL of the form <a target="_blank" href="https://abc123.ngrok.io">https://abc123.ngrok.io</a></p>
<p>Which is pointing to port 5000 of your system.</p>
<p><strong>Note:</strong> <strong><em>We have not setup a server running on port 5000 yet</em></strong>, but that is fine since ngrok is independent of that, and exposes the port regardless.</p>
<h3 id="heading-creating-the-slack-application"><strong>Creating the Slack Application</strong></h3>
<ol>
<li><p>Go to Slack API Apps – <a target="_blank" href="https://api.slack.com/apps">https://api.slack.com/apps</a></p>
</li>
<li><p>Click on Create App, and select the option ‘From a manifest’</p>
</li>
<li><p>Copy the <a target="_blank" href="https://github.com/DrDroidLab/slack-ai-bot-builder/blob/main/slack_manifest.json">manifest json</a> from the repository and <strong><mark>replace the placeholder ‘&lt;hostname&gt;’ with the HTTPS URL you got from ngrok. Include the “https://” part as well.</mark></strong></p>
</li>
<li><p>Install the app in your workspace.</p>
</li>
<li><p>Copy the credentials for the slack application into credentials.yaml</p>
<ol>
<li><p>app_id, app_name and signing secret can be found in the basic information tab.</p>
</li>
<li><p>Bot-auth-token can be found in the Oauth &amp; Permissions tab.</p>
</li>
<li><p>openai_key from openai - if you plan to use AI based workflows.</p>
</li>
</ol>
</li>
<li><p>Create a channel called #drdroid-slack-bot-tester in your slack workspace &amp; add the bot to the channel.</p>
</li>
</ol>
<h3 id="heading-setting-up-the-bot-server"><strong>Setting up the bot server</strong></h3>
<p>Run the following commands to set up your virtual environment and activate it.</p>
<pre><code class="lang-plaintext">uv env venv
source .venv/bin/activate
</code></pre>
<p>Install dependencies using:</p>
<pre><code class="lang-plaintext">uv sync
</code></pre>
<p>Now we can finally run the bot server using:</p>
<pre><code class="lang-plaintext">uv run python app.py
</code></pre>
<p>The server is now running on port 5000, and exposed to the outside world via your ngrok tunnel.</p>
<h3 id="heading-testing-the-bot"><strong>Testing the bot</strong></h3>
<p>Add the bot to the drdroid-slack-bot-tester workspace that you had previously created.</p>
<p>And just type in a ‘hi’. The bot should send you a sample response.</p>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1753596133455/a92590b2-6b31-479b-a209-015fa52872b3.png" alt class="image--center mx-auto" /></p>
<h2 id="heading-step-2-intehttpsgithubcomdrdroidlabslack-ai-bot-buildergrating-aihttpsngrokcomdocsgetting-started"><a target="_blank" href="https://github.com/DrDroidLab/slack-ai-bot-builder"><strong>Step 2: Inte</strong></a><a target="_blank" href="https://ngrok.com/docs/getting-started/"><strong>grating AI</strong></a></h2>
<p>There is already an example workflow for AI (name: "chatbot") in the <a target="_blank" href="https://github.com/DrDroidLab/slack-ai-bot-builder/blob/main/workflows.yaml">workflows.yaml</a>. It is just boilerplate code, you can modify it as required. For now, you can tag your bot in the drdroid-slack-bot-tester, and chat with it.</p>
<ul>
<li>Add your OpenAI/LLM key</li>
</ul>
<p>For example:</p>
<pre><code class="lang-plaintext">Message in Slack: chatbot How to debug Kubernetes CrashLoopBackOff error? @bot
Message in Slack: chatbot I'm getting this alert. What does it mean? @bot
</code></pre>
<h2 id="heading-step-3-making-the-bot-an-agent-by-giving-ai-access-to-different-tools-grafana-for-demo"><strong>Step 3: Making the bot an Agent by giving AI access to different tools (Grafana for demo)</strong></h2>
<p>MCP Servers help abstract out any API and making them accessible to AI.</p>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1753595526326/211ca3ee-f4b8-45f2-a1f0-3520850160a6.png" alt class="image--center mx-auto" /></p>
<h3 id="heading-setting-up-grafana-mcp-server"><strong>Setting up Grafana MCP Server</strong></h3>
<p>Clone the following repository:</p>
<p>Repository URL - <a target="_blank" href="https://github.com/DrDroidLab/grafana-mcp-server">https://github.com/DrDroidLab/grafana-mcp-server</a></p>
<p>Navigate into the root directory of the repository.</p>
<p><mark>Populate the </mark> <code>src/grafana_mcp_server/config.yaml</code> <mark> with your grafana credentials.</mark></p>
<p>Install and setup the dependencies using:</p>
<pre><code class="lang-plaintext">uv venv .venv
source .venv/bin/activate
uv sync
</code></pre>
<p>Run the MCP server:</p>
<pre><code class="lang-plaintext">uv run -m src.grafana_mcp_server.mcp_server
</code></pre>
<p>Your MCP server is now running on port 8000.</p>
<h3 id="heading-creating-an-ai-grafana-workflow-in-slack-bot-builder"><strong>Creating an AI Grafana workflow in slack-bot-builder</strong></h3>
<p>There is already an example workflow for grafana ai in the workflows.yaml. </p>
<p>We will be running the script scripts/grafana_ai_tool.py in this workflow, it is just boilerplate code, you can modify it as required.<br />Now you can tag your bot in the drdroid-slack-bot-tester, and ask it to do various things from grafana.<br />For example:</p>
<pre><code class="lang-plaintext">Message in Slack: Fetch me logs from the currencyservice in grafana ai. 
Message in Slack: Fetch and analyse the Go Microservices dashboard from grafana ai
</code></pre>
<h2 id="heading-next-steps">Next Steps:</h2>
<p>Now that you’ve been able to setup a bot, here are a few things you can do:</p>
<ul>
<li><p>Productionise it from your current ngrok setup to a static endpoint</p>
</li>
<li><p>Integrate with Grafana or open source <a target="_blank" href="https://glama.ai/mcp/servers">MCP servers</a> of your favourite tool that you want to leverage for automation</p>
</li>
<li><p>Add custom prompts and scripts</p>
</li>
</ul>
<p>Stuck anywhere? Ask on our <a target="_blank" href="https://discord.gg/AQ3tusPtZn">Discord</a></p>
]]></content:encoded></item><item><title><![CDATA[GitOps for Alerting: How to Manage Alert Rules Like Code]]></title><description><![CDATA[It's 2 AM. Production is on fire. You need to adjust an alert threshold that's been firing false positives all week.
You log into Grafana, click through three nested menus, find the alert, and bump the threshold from 80% to 85%. Crisis averted. You g...]]></description><link>https://notes.drdroid.io/gitops-for-alerting-how-to-manage-alert-rules-like-code</link><guid isPermaLink="true">https://notes.drdroid.io/gitops-for-alerting-how-to-manage-alert-rules-like-code</guid><category><![CDATA[#AIOps]]></category><dc:creator><![CDATA[SriNikitha Thummanapalli]]></dc:creator><pubDate>Fri, 18 Jul 2025 11:52:43 GMT</pubDate><enclosure url="https://cdn.hashnode.com/res/hashnode/image/upload/v1752833279010/4872f8bc-f608-469b-a39e-8fb172e66e6b.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>It's 2 AM. Production is on fire. You need to adjust an alert threshold that's been firing false positives all week.</p>
<p>You log into Grafana, click through three nested menus, find the alert, and bump the threshold from 80% to 85%. Crisis averted. You go back to bed.</p>
<p>Two weeks later, during a postmortem, someone asks: "Who changed the CPU alert threshold? And why?"</p>
<p>Silence. Nobody remembers. There's no history. No context. No way to know if this was a temporary hack or a deliberate tuning decision. Worse, when you refresh your staging environment, the old threshold returns because the change only lived in the production UI.</p>
<p>Sound familiar? You're not alone. This is how most teams manage alerts—and it's fundamentally broken.</p>
<h2 id="heading-why-managing-alerts-in-dashboards-doesnt-scale">Why Managing Alerts in Dashboards Doesn't Scale</h2>
<p>We've spent the last decade moving infrastructure to code. Terraform for cloud resources. Helm charts for Kubernetes. Ansible for configuration. Yet somehow, our alert rules—critical infrastructure that wakes up engineers—still live in UI dashboards like it's 2010.</p>
<p>The problems compound quickly:</p>
<p><strong>No version history</strong>: When did this alert last change? Who changed it? Why? Your Grafana dashboard shrugs.</p>
<p><strong>No peer review</strong>: A junior engineer can accidentally change a critical alert threshold with zero oversight. Try doing that with production code.</p>
<p><strong>No rollback capability</strong>: That "quick fix" that made things worse? Good luck remembering the old values.</p>
<p><strong>Environment drift</strong>: Production alerts diverge from staging. Dev environments have different rules. Chaos ensues.</p>
<p><strong>No ownership tracking</strong>: Who owns this alert? Which team should review changes? The UI doesn't care.</p>
<p>Your infrastructure evolves constantly. Services scale. Traffic patterns shift. Performance characteristics change. But alerts configured through dashboards remain frozen in time, slowly becoming less relevant until they're just noise.</p>
<p>Here's the thing: <strong>alert rules are infrastructure-as-code too</strong>. They define critical system behavior. They impact your team's quality of life. They deserve the same rigor as any other code.</p>
<p>Enter GitOps for alerts—where alert definitions live in version control, changes happen through pull requests, and every modification is tracked, reviewed, and reversible.</p>
<h2 id="heading-what-is-gitops-for-alerting">What is GitOps for Alerting?</h2>
<p>GitOps for alerting is beautifully simple: store your alert rules as code in Git, manage changes through pull requests, and deploy automatically. Just like any other infrastructure.</p>
<p>Most modern monitoring tools already support this:</p>
<ul>
<li><p><strong>Prometheus</strong>: Alert rules in YAML files</p>
</li>
<li><p><strong>Alertmanager</strong>: Routing configuration as code</p>
</li>
<li><p><strong>Grafana</strong>: Alerts exportable as JSON</p>
</li>
<li><p><strong>Datadog</strong>: Monitors manageable via Terraform</p>
</li>
<li><p><strong>New Relic</strong>: Alerts configurable through their API/Terraform</p>
</li>
</ul>
<p>Here's what a typical structure looks like:</p>
<pre><code class="lang-bash">/alerts/
  frontend-service.yaml
  database.yaml
  redis.yaml

/teams/
  payments/
    api-alerts.yaml
    database-alerts.yaml
  platform/
    infrastructure-alerts.yaml
    kubernetes-alerts.yaml
</code></pre>
<p>A Prometheus alert rule might look like:</p>
<pre><code class="lang-yaml"><span class="hljs-attr">groups:</span>
  <span class="hljs-bullet">-</span> <span class="hljs-attr">name:</span> <span class="hljs-string">frontend-service</span>
    <span class="hljs-attr">rules:</span>
      <span class="hljs-bullet">-</span> <span class="hljs-attr">alert:</span> <span class="hljs-string">HighErrorRate</span>
        <span class="hljs-attr">expr:</span> <span class="hljs-string">rate(http_requests_total{status=~"5.."}[5m])</span> <span class="hljs-string">&gt;</span> <span class="hljs-number">0.05</span>
        <span class="hljs-attr">for:</span> <span class="hljs-string">5m</span>
        <span class="hljs-attr">labels:</span>
          <span class="hljs-attr">severity:</span> <span class="hljs-string">warning</span>
          <span class="hljs-attr">service:</span> <span class="hljs-string">frontend</span>
          <span class="hljs-attr">team:</span> <span class="hljs-string">frontend-team</span>
        <span class="hljs-attr">annotations:</span>
          <span class="hljs-attr">summary:</span> <span class="hljs-string">"High error rate on <span class="hljs-template-variable">{{ $labels.instance }}</span>"</span>
          <span class="hljs-attr">description:</span> <span class="hljs-string">"Error rate is <span class="hljs-template-variable">{{ $value }}</span> (threshold 0.05)"</span>
          <span class="hljs-attr">runbook:</span> <span class="hljs-string">"https://wiki.company.com/runbooks/frontend-errors"</span>
          <span class="hljs-attr">owner:</span> <span class="hljs-string">"frontend-oncall@company.com"</span>
<span class="hljs-attr">The benefits are immediate:</span>
</code></pre>
<p>✅ <strong>Traceability</strong>: Every change is a commit. Git blame tells you who changed what and when.</p>
<p>✅ <strong>Peer review</strong>: Alert changes go through PR reviews. No more accidental 3 AM threshold adjustments.</p>
<p>✅ <strong>Consistency</strong>: Deploy the same alerts across all environments. No more production/staging drift.</p>
<p>✅ <strong>Rollback capability</strong>: Bad change? <code>git revert</code> and you're back to working alerts.</p>
<p>✅ <strong>Documentation</strong>: PR descriptions explain why changes were made. Context is preserved forever.</p>
<h2 id="heading-but-gitops-alone-isnt-enough">But GitOps Alone Isn't Enough</h2>
<p>Here's the plot twist: GitOps for alerts solves the <em>how</em> but not the <em>what</em>.</p>
<p>You now have beautiful, version-controlled alert rules. Every change is reviewed and tracked. But you still don't know <strong>which rules need updating</strong>.</p>
<p>Your Git repo becomes a graveyard of alert rules that might or might not be relevant:</p>
<ul>
<li><p>That CPU alert from 2019 when you ran on smaller instances</p>
</li>
<li><p>The memory warning tuned for your old Java app (you've since moved to Go)</p>
</li>
<li><p>The latency threshold set when you had 100 users (you now have 10,000)</p>
</li>
</ul>
<p>You've traded one problem for another. Instead of stale alerts in dashboards, you have stale alerts in Git. They're better organized, sure, but still noisy.</p>
<p>This is where most GitOps alerting stories end. Teams implement the framework but lack the feedback loop to keep it healthy. Alert rules accumulate like sediment. Engineers suffer in silence because "at least it's in Git now."</p>
<h2 id="heading-using-alert-insights-to-drive-gitops-changes">Using Alert Insights to Drive GitOps Changes</h2>
<h3 id="heading-let-real-alert-data-guide-your-pull-requests">Let real alert data guide your pull requests</h3>
<p>The missing piece is data. You need to know which alerts are actually problematic before you can fix them. This is where <strong>DrDroid's Alert Insights</strong> transforms GitOps from a theoretical improvement into a practical solution.</p>
<p>Alert Insights analyzes your live production alerts and tells you:</p>
<ul>
<li><p><strong>Which alerts fired most frequently last week</strong>: Your noisiest offenders, ranked</p>
</li>
<li><p><strong>Which alerts were ignored</strong>: Clear signal of rules that need removal</p>
</li>
<li><p><strong>Which alerts lack owners or runbooks</strong>: Quality issues to address</p>
</li>
<li><p><strong>Suggested changes</strong>: Specific recommendations to mute, tweak, or archive</p>
</li>
</ul>
<p>Now GitOps becomes powerful. You're not guessing which alert rules to update—you have data.</p>
<p>✅ <strong>Workflow Example:</strong></p>
<p><strong>Monday: Run Alert Insights</strong></p>
<p>`Top 3 Noisy Alerts:</p>
<ol>
<li><p>redis_memory_warning - 127 fires, 0 actions taken</p>
</li>
<li><p>api_latency_high - 89 fires, acknowledged but not investigated</p>
</li>
<li><p>cpu_usage_critical - 45 fires, all during deploy windows`</p>
</li>
</ol>
<p><strong>Tuesday: Create targeted PRs</strong></p>
<pre><code class="lang-bash">git checkout -b fix/reduce-redis-memory-noise
<span class="hljs-comment"># Edit alerts/redis.yaml</span>
<span class="hljs-comment"># Increase threshold from 70% to 80% based on actual usage patterns</span>
git commit -m <span class="hljs-string">"Increase Redis memory threshold to reduce false positives

Alert Insights showed 127 fires with 0 actions last week. Analysis shows Redis memory naturally spikes to 75% during cache warmup."</span>
</code></pre>
<p><strong>Wednesday: Review and merge</strong></p>
<ul>
<li><p>Team reviews the PR</p>
</li>
<li><p>Links to Alert Insights data provide context</p>
</li>
<li><p>Changes deploy automatically</p>
</li>
</ul>
<p><strong>Thursday: Validate impact</strong></p>
<ul>
<li><p>Alert noise drops immediately</p>
</li>
<li><p>Next week's Alert Insights confirms improvement</p>
</li>
</ul>
<p>The feedback loop is complete. You're not just organizing alerts better—you're systematically improving them based on real data.</p>
<p>➡️ <strong>🛠️ Want a GitOps-ready alert audit? 👉</strong> <a target="_blank" href="https://drdroid.io/doctor-droid-slack-integration"><strong>Run DrDroid's Alert Insights</strong></a> <strong>and get actionable suggestions in minutes.</strong></p>
<h2 id="heading-recommended-practices-for-gitops-alert-management">Recommended Practices for GitOps Alert Management</h2>
<h3 id="heading-use-clear-filenames-per-servicecomponent">🔍 Use clear filenames per service/component</h3>
<p>Don't create a monolithic <code>alerts.yaml</code>. Break rules into logical groups:</p>
<pre><code class="lang-plaintext">/alerts/
  services/
    payment-api.yaml
    user-service.yaml
  infrastructure/
    kubernetes-nodes.yaml
    database-cluster.yaml
  business/
    checkout-flow.yaml
    user-engagement.yaml
</code></pre>
<h3 id="heading-add-labelstags-to-help-alert-insights-map-alerts-to-owners">🔄 Add labels/tags to help Alert Insights map alerts to owners</h3>
<p>Every alert should include:</p>
<p>yaml</p>
<p><code>labels: team: payments service: payment-api environment: production severity: P2</code></p>
<p>This metadata powers Alert Insights' analysis and recommendations.</p>
<h3 id="heading-validate-rules-with-test-alerts-in-staging">🧪 Validate rules with test alerts in staging</h3>
<p>Before merging, trigger test conditions:</p>
<pre><code class="lang-bash"><span class="hljs-comment"># Simulate high error rate</span>
curl -X POST http://prometheus:9090/api/v1/series \
-d <span class="hljs-string">'match[]=up{job="frontend"}'</span>
🔁 Link PRs to weekly alert review
</code></pre>
<h2 id="heading-common-gitops-pitfalls-to-avoid">Common GitOps Pitfalls to Avoid</h2>
<h3 id="heading-bulk-silencing-alerts-without-context">❌ Bulk silencing alerts without context</h3>
<p>"Let's just comment out all the noisy alerts" is tempting but dangerous. Use Alert Insights to understand <em>why</em> alerts are noisy before acting.</p>
<h3 id="heading-committing-rules-without-reviews">❌ Committing rules without reviews</h3>
<p>The whole point of GitOps is peer review. Don't bypass it with direct commits, even for "quick fixes."</p>
<h3 id="heading-no-tagging-alert-insights-cant-map-alerts-to-services">❌ No tagging = Alert Insights can't map alerts to services</h3>
<p>Without proper labels, you lose the ability to analyze alerts by team, service, or severity. Enforce tagging standards.</p>
<h3 id="heading-alert-rules-diverging-across-environments">❌ Alert rules diverging across environments</h3>
<p>Use templating to keep staging and production alerts synchronized:</p>
<pre><code class="lang-yaml"><span class="hljs-comment"># values-prod.yaml</span>
<span class="hljs-attr">cpu_threshold:</span> <span class="hljs-number">80</span>
<span class="hljs-attr">memory_threshold:</span> <span class="hljs-number">85</span>

<span class="hljs-comment"># values-staging.yaml</span>
<span class="hljs-attr">cpu_threshold:</span> <span class="hljs-number">90</span> <span class="hljs-comment"># Higher tolerance in staging</span>
<span class="hljs-attr">memory_threshold:</span> <span class="hljs-number">90</span>
</code></pre>
<h2 id="heading-final-take-let-data-drive-your-alert-rule-changes">Final Take — Let Data Drive Your Alert Rule Changes</h2>
<p>GitOps gives you the framework for managing alerts professionally. Version control, peer review, and rollback capabilities bring alerts into the modern era.</p>
<p>But framework without data is just organized chaos. You need to know which alerts to fix, how to fix them, and whether your fixes worked.</p>
<p>Alert Insights provides that missing data layer. It tells you which alert rules are hurting your team, suggests specific improvements, and validates that your changes actually reduced noise.</p>
<p>Together, they create a powerful feedback loop:</p>
<ol>
<li><p>Alert Insights identifies problematic alerts</p>
</li>
<li><p>GitOps enables reviewed, tracked changes</p>
</li>
<li><p>Automated deployment ensures consistency</p>
</li>
<li><p>Next week's Alert Insights validates improvement</p>
</li>
</ol>
<p>This isn't theoretical. Teams using this approach report 50-70% reduction in alert noise within weeks. On-call engineers sleep better. Real incidents get proper attention. Alert quality becomes a measurable, improvable metric.</p>
<p>Your alerts deserve the same engineering rigor as your code. GitOps provides the foundation. Alert Insights provides the intelligence. Together, they transform alerting from a necessary evil into a competitive advantage.</p>
<p>➡️ <strong>✍️ Want to make smarter, reviewable changes to your alerts? 👉</strong> <a target="_blank" href="https://aiops.drdroid.io/">Run AIOps</a> <strong>and let your alerts tell you what to fix.</strong></p>
]]></content:encoded></item><item><title><![CDATA[3 Tools That Help Reduce Alert Fatigue (With Trade-offs)]]></title><description><![CDATA[We live in the age of "vibecoding."
Your engineers ship features at lightning speed. AI copilots autocomplete entire functions. CI/CD pipelines deploy to production in minutes. Modern development has become a symphony of efficiency, with developers o...]]></description><link>https://notes.drdroid.io/3-tools-that-help-reduce-alert-fatigue-with-trade-offs</link><guid isPermaLink="true">https://notes.drdroid.io/3-tools-that-help-reduce-alert-fatigue-with-trade-offs</guid><category><![CDATA[alert noise]]></category><category><![CDATA[alert-insights]]></category><dc:creator><![CDATA[SriNikitha Thummanapalli]]></dc:creator><pubDate>Fri, 18 Jul 2025 11:52:02 GMT</pubDate><enclosure url="https://cdn.hashnode.com/res/hashnode/image/upload/v1752831431954/cb658a85-d1de-48db-b2e2-c3de9f9ebeff.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>We live in the age of "vibecoding."</p>
<p>Your engineers ship features at lightning speed. AI copilots autocomplete entire functions. CI/CD pipelines deploy to production in minutes. Modern development has become a symphony of efficiency, with developers operating at 10x the speed of just five years ago.</p>
<p>But there's one part of your stack that's stuck in 2010: your alerts.</p>
<p>While your team vibecodes their way through complex distributed systems, your alerting engine still screams about every CPU spike, memory blip, and network hiccup like it's the apocalypse. It's like having a Ferrari engine attached to horse-and-buggy wheels. The cognitive dissonance is jarring—and it's killing your team's productivity.</p>
<h2 id="heading-why-alert-fatigue-is-a-real-problem-in-2025"><strong>Why Alert Fatigue Is a Real Problem in 2025</strong></h2>
<p>Here's the absurd reality: The same engineer who just deployed a sophisticated ML model in production gets woken up at 3 AM because a health check endpoint took 501ms instead of 500ms to respond. The developer who elegantly orchestrated a microservices migration gets paged because a pod restarted—something Kubernetes is literally designed to do automatically.</p>
<p>Modern infrastructure has exploded in complexity. You're running hundreds of microservices, each generating alerts. Kubernetes adds its own layer of notifications. Cloud providers, APMs, and security tools all want their voice heard. The result? An endless stream of "urgent" notifications flooding Slack channels and PagerDuty rotations.</p>
<p>But unlike your codebase—which has intelligent linters, smart IDEs, and AI-powered suggestions—your alerts remain dumb. They can't distinguish between:</p>
<ul>
<li><p>A temporary spike during garbage collection vs. a memory leak</p>
</li>
<li><p>A planned scaling event vs. an unexpected traffic surge</p>
</li>
<li><p>A self-healing Kubernetes pod restart vs. a critical service failure</p>
</li>
</ul>
<p>The real problem: <strong>You don't know which alerts matter anymore.</strong></p>
<p>Your engineers have adapted the only way they can—by tuning out. When every alert claims to be critical but most are noise, even genuine emergencies get ignored. It's the monitoring equivalent of crying wolf, except the wolf is paging your on-call engineer every 30 minutes.</p>
<p>What you need aren't more dashboards visualizing the chaos. You need intelligent tools that understand context, learn patterns, and <strong>show you what's noisy and help you take action</strong>. Let's examine three approaches to bringing your alerts into the modern era.</p>
<h2 id="heading-tool-1-drdroid"><strong>Tool #1 – DrDroid</strong></h2>
<h3 id="heading-best-for-real-time-visibility-into-noisy-alerts-across-any-stack"><strong>Best for: Real-time visibility into noisy alerts, across any stack</strong></h3>
<p>DrDroid represents the first generation of truly intelligent alerting tools. While your engineers use AI to write code faster, DrDroid uses intelligence to make your alerts smarter.</p>
<p>The platform integrates with your existing stack—Slack, Prometheus, New Relic, OpenTelemetry, and more. But what sets it apart is the <strong>Alert Insights</strong> feature, which applies actual intelligence to your alert patterns:</p>
<ul>
<li><p><strong>Which alerts are flapping?</strong> Just like a smart IDE highlights code smells, DrDroid identifies alerts that repeatedly fire and resolve—clear indicators of misconfiguration.</p>
</li>
<li><p><strong>Which alerts are being ignored?</strong> By analyzing engineer behavior, it spots alerts that get dismissed without action. If developers ignore an alert 100% of the time, why is it still paging them?</p>
</li>
<li><p><strong>Which alerts lack runbooks or clear owners?</strong> Nothing frustrates a vibecoding engineer more than context-switching to an alert with zero information about what to do.</p>
</li>
</ul>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1752754957104/e354b06d-d477-479d-939c-4c03a3338299.png" alt /></p>
<p>DrDroid doesn't just identify problems—it suggests fixes:</p>
<ul>
<li><p>Automatically mute alerts during deployment windows</p>
</li>
<li><p>Disable alerts that have never correlated with customer impact</p>
</li>
<li><p>Add intelligent conditions (like requiring sustained threshold breaches)</p>
</li>
<li><p>Enrich alerts with missing context, runbooks, and correlation data</p>
</li>
</ul>
<p>The platform's <strong>auto-debugging</strong> capabilities are particularly impressive. When an alert fires, DrDroid automatically pulls relevant logs, metrics, traces, and even recent code changes. It's like having an AI copilot for incident response.</p>
<p>Consider this scenario: Your payment service alerts on high latency every day at 2 PM. DrDroid notices the pattern, correlates it with a scheduled batch job, and suggests either suppressing the alert during that window or adjusting the threshold. What took hours of manual analysis now happens automatically.</p>
<h3 id="heading-trade-offs"><strong>Trade-offs</strong></h3>
<p>DrDroid is built for modern, Slack-first teams. If your organization has traditional processes requiring all alerts to flow through legacy ITSM tools, adoption might face resistance. As a newer platform, some enterprise compliance features are still maturing.</p>
<p>➡️ <strong>🧠 Want to know which alerts your team should disable? 👉</strong> <a target="_blank" href="https://drdroid.io/doctor-droid-slack-integration"><strong>Explore DrDroid's Alert Insights</strong></a> <strong>— loved by SREs to reduce alert fatigue.</strong></p>
<h2 id="heading-tool-2-bigpanda"><strong>Tool #2 – BigPanda</strong></h2>
<h3 id="heading-best-for-enterprise-scale-alert-correlation"><strong>Best for: Enterprise-scale alert correlation</strong></h3>
<p>BigPanda takes a different approach—using machine learning to group related alerts into incidents. When a database issue triggers alerts across 20 services, BigPanda recognizes the pattern and presents them as one incident.</p>
<p>For large enterprises with complex systems, this correlation can help. The platform learns relationships between components and can reduce the number of incidents operators review. It also integrates deeply with enterprise tools like ServiceNow and Dynatrace.</p>
<h3 id="heading-trade-offs-1"><strong>Trade-offs</strong></h3>
<p>Here's where the contrast with modern development becomes stark. While your engineers deploy code in minutes, BigPanda requires <strong>months of setup</strong>. While developers use intuitive tools that work out-of-the-box, BigPanda demands extensive metadata configuration and alert standardization.</p>
<p>More critically, BigPanda doesn't make individual alerts smarter—it just groups dumb alerts better. Those flapping alerts your engineers hate? Still firing, just bundled together. It's like organizing spam into folders instead of fixing your spam filter.</p>
<p>The platform is also expensive, often requiring dedicated administrators and cross-team coordination. For teams used to the speed of modern development, BigPanda's implementation timeline feels like stepping back in time.</p>
<h2 id="heading-tool-3-pagerduty-analytics"><strong>Tool #3 – PagerDuty Analytics</strong></h2>
<h3 id="heading-best-for-trend-visibility-inside-the-pagerduty-ecosystem"><strong>Best for: Trend visibility inside the PagerDuty ecosystem</strong></h3>
<p>PagerDuty Analytics provides retrospective dashboards showing alert volume, MTTR, and on-call load. For teams already using PagerDuty, it offers visibility into historical patterns and trends.</p>
<p>The analytics can be useful for quarterly reviews and capacity planning. You can see which services generate the most alerts and track improvements over time.</p>
<h3 id="heading-trade-offs-2"><strong>Trade-offs</strong></h3>
<p>The limitations mirror the gap between modern development and legacy monitoring. While your engineers get real-time feedback from their tools, PagerDuty Analytics is <strong>retrospective only</strong>. It tells you that Service X generated 500 alerts last month but not which ones were false positives or what to do about them.</p>
<p>It only analyzes alerts flowing through PagerDuty, missing Slack notifications and other channels. The insights are descriptive, not prescriptive—you see the problem visualized but get no help fixing it. And it requires expensive premium tiers, adding cost without adding intelligence.</p>
<h2 id="heading-comparison-table"><strong>Comparison Table</strong></h2>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1752754340622/c0dfb6f6-7cef-478d-b60b-ccfbbaa773ba.png" alt /></p>
<h2 id="heading-final-thoughts-your-alerts-should-be-as-smart-as-your-code"><strong>Final Thoughts — Your Alerts Should Be as Smart as Your Code</strong></h2>
<p>We've entered an era where engineers can literally describe what they want to build and watch AI generate the code. They deploy with confidence, iterate rapidly, and ship features that would have taken months in mere days.</p>
<p>Yet these same engineers—these 10x vibecoding machines—are still being interrupted by alerts that would have been considered noisy a decade ago.</p>
<p>The disconnect is unsustainable. You can't run a modern engineering organization with stone-age alerting. Your monitoring needs to evolve to match the sophistication of your development practices.</p>
<p>BigPanda and PagerDuty show you the problem in high resolution. <strong>Only DrDroid's Alert Insights actually makes your alerts smarter</strong>—identifying what's broken, why it's noisy, and exactly how to fix it.</p>
<p>The future of monitoring isn't better dashboards or fancier grouping algorithms. It's intelligent systems that understand context, learn from patterns, and proactively help you maintain signal-to-noise ratio. It's alerts that are as smart as the engineers they're interrupting.</p>
<p>Your team deserves alerting infrastructure that matches their development velocity. Stop letting 2010-era alerts slow down your 2025 engineering team.</p>
<p>➡️ <strong>💡 Ready to reduce alert fatigue the smart way? 👉</strong> <a target="_blank" href="https://drdroid.io/doctor-droid-slack-integration"><strong>Start using Alert Insights</strong></a> <strong>to find and fix noisy alerts today — no config needed.</strong></p>
]]></content:encoded></item><item><title><![CDATA[A Practical Framework to Reduce Alert Noise (Without Missing Incidents)]]></title><description><![CDATA[Every SRE has been there.
Fed up with alert fatigue, you go on a muting spree. That flaky health check? Silenced. The CPU warning that fires during deploys? Disabled. The memory alert that triggers during garbage collection? Gone.
For a blissful week...]]></description><link>https://notes.drdroid.io/a-practical-framework-to-reduce-alert-noise-without-missing-incidents</link><guid isPermaLink="true">https://notes.drdroid.io/a-practical-framework-to-reduce-alert-noise-without-missing-incidents</guid><category><![CDATA[alert-insights]]></category><category><![CDATA[monitoring]]></category><category><![CDATA[AI]]></category><category><![CDATA[alerting]]></category><dc:creator><![CDATA[SriNikitha Thummanapalli]]></dc:creator><pubDate>Fri, 18 Jul 2025 10:23:21 GMT</pubDate><enclosure url="https://cdn.hashnode.com/res/hashnode/image/upload/v1752831592236/dc329375-d56a-4043-b40e-4c1815e6bf86.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Every SRE has been there.</p>
<p>Fed up with alert fatigue, you go on a muting spree. That flaky health check? Silenced. The CPU warning that fires during deploys? Disabled. The memory alert that triggers during garbage collection? Gone.</p>
<p>For a blissful week, your on-call rotation is peaceful. Engineers are sleeping through the night. Slack channels are quiet. Life is good.</p>
<p>Then it happens. A real incident slips through. Customer complaints pour in. Your CEO wants answers. And suddenly, those "noisy" alerts you disabled don't seem so unnecessary anymore.</p>
<p>Here's the uncomfortable truth: anyone can reduce alert noise by turning off alerts. The real challenge—the one that separates good SRE teams from great ones—is reducing noise <strong>without sacrificing coverage</strong>.</p>
<h2 id="heading-why-reducing-alert-noise-is-harder-than-it-sounds">Why Reducing Alert Noise Is Harder Than It Sounds</h2>
<p>The naive approach to alert fatigue is seductively simple: just turn off the annoying alerts. But this creates a dangerous blind spot. That CPU alert might be noisy 99% of the time, but what about the 1% when it signals a real problem?</p>
<p>The opposite extreme isn't better. Some teams, burned by missed incidents, keep every alert active "just in case." They end up with hundreds of alerts that cry wolf, training engineers to ignore everything—including real emergencies.</p>
<p>The solution isn't choosing between noise and coverage. It's building a systematic approach that maintains visibility while eliminating false positives. High-performing SRE teams follow a <strong>4-phase framework</strong> that transforms chaotic alerting into intelligent monitoring.</p>
<p>This framework isn't theoretical—it's battle-tested by teams managing hundreds of services in production. And with modern tools like <strong>Alert Insights</strong>, you can measure and validate your improvements with data, not guesswork.</p>
<p>Let's dive into each phase.</p>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1752831961901/f3951003-f149-4ffa-90a0-b9bad0526240.png" alt class="image--center mx-auto" /></p>
<h2 id="heading-phase-1-start-with-coverage-not-silence">Phase 1 – Start with Coverage, Not Silence</h2>
<h3 id="heading-the-mistake-most-teams-make-start-muting">The mistake most teams make: start muting</h3>
<p>When alert fatigue hits, the instinctive response is to start silencing alerts. It feels productive—each muted alert is one less interruption. But this approach is backwards.</p>
<p>Before you disable a single alert, you need to understand what you're actually trying to monitor. This means mapping your core Service Level Objectives (SLOs) and Service Level Indicators (SLIs) to your alerting strategy.</p>
<p>Most teams rely too heavily on infrastructure alerts—CPU usage, memory consumption, disk space. These are important, but they're indirect signals. A service can have high CPU usage while serving customers perfectly. Conversely, it can have normal resource usage while completely failing its primary function.</p>
<p>Instead, start with user-facing signals:</p>
<ul>
<li><p><strong>Failed user logins</strong> (not just authentication service uptime)</p>
</li>
<li><p><strong>Checkout completion rates</strong> (not just payment gateway availability)</p>
</li>
<li><p><strong>API response times at the 95th percentile</strong> (not just average latency)</p>
</li>
<li><p><strong>Database query failures</strong> (not just connection pool metrics)</p>
</li>
</ul>
<p>Map these business-critical indicators first. Only after you have comprehensive coverage of what matters should you start tuning what doesn't.</p>
<p>✅ <strong>Principle:</strong> Only tune alerts after coverage is solid. It's better to have noisy but comprehensive alerting than quiet but blind monitoring.</p>
<h2 id="heading-phase-2-assign-ownership">Phase 2 – Assign Ownership</h2>
<h3 id="heading-every-alert-should-have-an-owner-a-service-and-a-runbook">Every alert should have an owner, a service, and a runbook</h3>
<p>Here's a dirty secret of most alerting systems: nobody owns the alerts. They fire into shared channels where responsibility diffuses across the team. When everyone is responsible, no one is accountable.</p>
<p>This shared ownership model is why alerts never improve. The payment team ignores database alerts because "that's infrastructure's problem." The infrastructure team ignores API latency alerts because "that's the app team's issue." Meanwhile, both alerts keep firing, and your on-call engineer suffers.</p>
<p>The fix is radical but simple: <strong>every alert must have a single owner</strong>. Not a team, not a rotation—a specific service and the team that owns it. This means:</p>
<ul>
<li><p>No more #alerts-general channels where everything dumps</p>
</li>
<li><p>No more "infrastructure noise" channels that everyone mutes</p>
</li>
<li><p>Each team gets their own alert destinations</p>
</li>
<li><p>Each team is accountable for their signal-to-noise ratio</p>
</li>
</ul>
<p>Implement this with proper tagging:</p>
<p><code>alert: HighAPILatency service: payment-api team: payments owner: payments-team@company.com escalation: payments-oncall severity: P2</code></p>
<p>When alerts have clear ownership, magic happens. The payments team suddenly cares about that flapping API alert because it's waking them up, not some random SRE. They'll fix it, tune it, or justify why it needs to stay.</p>
<p>✅ <strong>Tip:</strong> Alerts without owners almost never get fixed. They become background noise that everyone learns to ignore.</p>
<h2 id="heading-phase-3-enrich-then-tune">Phase 3 – Enrich, Then Tune</h2>
<h3 id="heading-rich-alerts-less-cognitive-load-faster-response">Rich alerts = less cognitive load = faster response</h3>
<p>Now that you have coverage and ownership, it's time to make your alerts actually useful. A bare-bones "Service X is down" notification forces engineers to context-switch, investigate, and piece together what's happening. Rich alerts provide everything upfront.</p>
<p>Essential enrichment includes:</p>
<ul>
<li><p><strong>Runbook links</strong>: Step-by-step remediation instructions</p>
</li>
<li><p><strong>Severity levels</strong>: Is this customer-impacting or internal-only?</p>
</li>
<li><p><strong>Business impact</strong>: How many users affected? Which features degraded?</p>
</li>
<li><p><strong>Recent changes</strong>: Did a deployment just go out?</p>
</li>
<li><p><strong>Historical context</strong>: Has this happened before? How was it fixed?</p>
</li>
</ul>
<p>But richness isn't verbosity. Don't dump entire log files into alerts. Instead, provide precisely what's needed for rapid decision-making.</p>
<p>Only after enrichment should you start tuning:</p>
<p><strong>Add intelligent conditions</strong>: Instead of alerting on every spike, require sustained problems:</p>
<ul>
<li><p>Alert only after 3 consecutive failures</p>
</li>
<li><p>Require issues to persist for 5 minutes</p>
</li>
<li><p>Use percentage-based thresholds (5% of requests failing vs. 10 absolute failures)</p>
</li>
</ul>
<p><strong>Adjust thresholds based on reality</strong>: That 80% CPU alert made sense with your old infrastructure. But if your auto-scaling kicks in at 70%, you're alerting on normal operations.</p>
<p><strong>Add flapping protection</strong>: If an alert fires and resolves repeatedly, it needs damping:</p>
<ul>
<li><p>Require state changes to persist before alerting</p>
</li>
<li><p>Group rapid-fire alerts into single notifications</p>
</li>
<li><p>Add cooldown periods between alerts</p>
</li>
</ul>
<p>✅ <strong>Insight:</strong> Context beats volume every time. One well-enriched alert is worth ten noisy notifications.</p>
<h2 id="heading-phase-4-use-data-to-improve-over-time">Phase 4 – Use Data to Improve Over Time</h2>
<h3 id="heading-enter-alert-insights-by-drdroid">Enter: Alert Insights by DrDroid</h3>
<p>Here's where most frameworks fail: they're static. Teams implement phases 1-3, declare victory, and move on. Six months later, they're back to alert fatigue because systems evolve but alerts don't.</p>
<p>You need a continuous feedback loop—a way to measure what's working and what's still broken. This is where <strong>Alert Insights</strong> becomes your secret weapon.</p>
<p>After implementing your alert structure, Alert Insights provides ongoing intelligence:</p>
<ul>
<li><p><strong>Which alerts are firing too often?</strong> That P1 alert that fires 50 times per week probably needs adjustment</p>
</li>
<li><p><strong>Which ones are being ignored?</strong> If engineers acknowledge but never act on an alert, it's pure noise</p>
</li>
<li><p><strong>Which lack runbooks or clear owners?</strong> Gaps in your enrichment strategy become visible</p>
</li>
<li><p><strong>What can be safely muted, disabled, or improved?</strong> Data-driven recommendations, not guesswork</p>
</li>
</ul>
<p>The workflow becomes systematic:</p>
<p><strong>Every sprint:</strong></p>
<ol>
<li><p>Review Alert Insights dashboard</p>
</li>
<li><p>Identify the top 3 worst offenders</p>
</li>
<li><p>Fix ownership, enrichment, or tuning for those alerts</p>
</li>
<li><p>Validate improvements in the next sprint</p>
</li>
<li><p>Repeat</p>
</li>
</ol>
<p>This creates a virtuous cycle. Your alerts get better every sprint. Your on-call experience improves measurably. And you maintain coverage while reducing noise.</p>
<p>➡️ <strong>🧠 Want a clear report on which alerts are hurting your team? 👉</strong> <a target="_blank" href="https://drdroid.io/doctor-droid-slack-integration"><strong>Run DrDroid Alert Insights</strong></a> <strong>— no config required.</strong></p>
<h2 id="heading-bringing-it-all-together-your-teams-framework">Bringing It All Together — Your Team's Framework</h2>
<p>Here's your systematic approach to intelligent alerting:</p>
<div class="hn-table">
<table>
<thead>
<tr>
<td><strong>Phase</strong></td><td><strong>Goal</strong></td><td><strong>Key Action</strong></td></tr>
</thead>
<tbody>
<tr>
<td>1. Coverage First</td><td>Avoid blind spots</td><td>Map alerts to SLOs</td></tr>
<tr>
<td>2. Ownership</td><td>Accountability</td><td>Assign alerts to teams</td></tr>
<tr>
<td>3. Enrichment &amp; Tuning</td><td>Faster resolution</td><td>Add context, reduce flapping</td></tr>
<tr>
<td>4. Feedback Loop</td><td>Continuous improvement</td><td>Use Alert Insights regularly</td></tr>
</tbody>
</table>
</div><p>This isn't a one-time project—it's an ongoing practice. Just like you continuously refactor code, you need to continuously refine alerts. The difference is that now you have a framework and the data to guide your decisions.</p>
<h2 id="heading-final-thought-you-cant-fix-what-you-dont-see">Final Thought — You Can't Fix What You Don't See</h2>
<p>Most teams exist in one of two failure modes. They either suffer in silence with alert fatigue, accepting it as the cost of observability. Or they oversimplify their alerting, creating dangerous blind spots that only become visible during incidents.</p>
<p>Real success looks different: high signal, low noise, and fast resolution. It's alerts that wake you up only when customer impact is imminent. It's notifications that include everything needed to respond. It's a system that improves continuously based on data, not opinions.</p>
<p>This framework gives you the path. Phase by phase, you can transform your alerting from a source of frustration into a competitive advantage. But frameworks only work when you can measure their impact.</p>
<p>Let <strong>Alert Insights</strong> be your guide. It shows what's working, what's broken, and exactly how to improve. No more guessing which alerts to tune. No more hoping you haven't created blind spots. Just data-driven improvements that make your team's life better.</p>
<p>Your engineers deserve better than alert fatigue. Your customers deserve better than missed incidents. This framework delivers both.</p>
<p>➡️ <strong>🛠️ Tired of guessing which alerts are noisy? 👉</strong> <a target="_blank" href="https://drdroid.io/doctor-droid-slack-integration"><strong>Try Alert Insights</strong></a> <strong>and start tuning your alerts based on real data.</strong></p>
]]></content:encoded></item><item><title><![CDATA[KubeCon + CloudNativeCon Europe 2025 Guide – London]]></title><description><![CDATA[Doctor Droid’s Guide to KubeCon + CloudNativeCon Europe 2025
Welcome to our complete guide for navigating KubeCon + CloudNativeCon Europe 2025 in London, England, running from 1–4 April 2025. Whether you’re a seasoned cloud native pro or new to the K...]]></description><link>https://notes.drdroid.io/kubecon-cloudnativecon-europe-2025-guide-london</link><guid isPermaLink="true">https://notes.drdroid.io/kubecon-cloudnativecon-europe-2025-guide-london</guid><category><![CDATA[KubeConLondon]]></category><category><![CDATA[Kubecon]]></category><category><![CDATA[Kubernetes]]></category><dc:creator><![CDATA[Jayesh Sadhwani]]></dc:creator><pubDate>Fri, 14 Feb 2025 08:32:01 GMT</pubDate><enclosure url="https://cdn.hashnode.com/res/hashnode/image/upload/v1739516728159/d72bee5d-8e33-4e45-b305-4bae01e37d47.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h2 id="heading-doctor-droids-guide-to-kubecon-cloudnativecon-europe-2025">Doctor Droid’s Guide to KubeCon + CloudNativeCon Europe 2025</h2>
<p>Welcome to our complete guide for navigating KubeCon + CloudNativeCon Europe 2025 in London, England, running from <strong>1–4 April 2025</strong>. Whether you’re a seasoned cloud native pro or new to the Kubernetes world, this guide has everything you need to maximize your conference experience. And don’t forget – visit the Doctor Droid booth at the event! Show us this blog to score an exclusive <strong>20% discount on your ticket</strong> plus a chance to receive special Doctor Droid credits!</p>
<hr />
<h2 id="heading-overview">Overview</h2>
<p>KubeCon + CloudNativeCon is the premier conference for Kubernetes, cloud-native technologies, and open source innovations. Organized by the Cloud Native Computing Foundation (CNCF), this flagship event gathers thousands of developers, engineers, and industry leaders to share ideas, network, and explore the latest trends shaping the future of cloud computing. In the heart of London, expect inspiring keynotes, deep-dive sessions, and engaging co-located events that cover everything from AI/ML integration to edge computing and beyond.</p>
<hr />
<h2 id="heading-access-types">Access Types</h2>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1739516645338/d12f9286-4b56-49db-86fa-fbd8279088b2.png" alt class="image--center mx-auto" /></p>
<p>KubeCon + CloudNativeCon Europe 2025 offers a single, all-inclusive pass that grants access to:</p>
<ul>
<li><p>All keynote sessions, breakout tracks, and panel discussions</p>
</li>
<li><p>Hands-on labs and workshops for a practical dive into cloud-native solutions</p>
</li>
<li><p>Co-located events hosted by CNCF and industry partners</p>
</li>
</ul>
<p>Additionally, there are special discounted options available for students and academic participants. For more details, check out the official <a target="_blank" href="https://events.linuxfoundation.org/kubecon-cloudnativecon-europe/">KubeCon + CloudNativeCon Europe website</a> for registration options.</p>
<hr />
<h2 id="heading-exclusive-kubecon-cloudnativecon-europe-2025-discount-save-20-on-tickets-courtesy-of-doctor-droid">Exclusive KubeCon + CloudNativeCon Europe 2025 Discount – Save 20% on Tickets, Courtesy of Doctor Droid!</h2>
<p>That’s right – Doctor Droid is proud to sponsor KubeCon + CloudNativeCon Europe 2025, and we’re offering you an exclusive 20% discount on your ticket! Here’s how to claim your savings:</p>
<ol>
<li><p><strong>Fill out our quick</strong> <a target="_blank" href="https://forms.gle/rMg1xAP34rA1jdM99"><strong>Google Form</strong></a> with your basic information.</p>
</li>
<li><p><strong>Receive your discount code</strong> directly in your inbox.</p>
</li>
<li><p><strong>Register</strong> on the official event website and enjoy your 20% savings!</p>
</li>
</ol>
<p>Hurry up – secure your discount today and get ready for an unforgettable experience in London!</p>
<hr />
<h2 id="heading-speakers-amp-tracks-at-kubecon-cloudnativecon-europe-2025">Speakers &amp; Tracks at KubeCon + CloudNativeCon Europe 2025</h2>
<p>The conference features an impressive line-up of industry thought leaders and technical experts across multiple tracks, including:</p>
<ul>
<li><p><strong>Kubernetes Operations:</strong> Best practices in deployment, scaling, and security.</p>
</li>
<li><p><strong>Cloud Security:</strong> Deep dives into safeguarding cloud native environments.</p>
</li>
<li><p><strong>AI &amp; Machine Learning:</strong> Innovations transforming how we manage and operate Kubernetes.</p>
</li>
<li><p><strong>Edge Computing:</strong> Exploring the future of distributed computing in real-world scenarios.</p>
</li>
</ul>
<p>Be sure to check out the detailed schedule on the <a target="_blank" href="https://events.linuxfoundation.org/kubecon-cloudnativecon-europe/">official event page</a> for a complete list of sessions and speakers.</p>
<hr />
<h2 id="heading-planning-your-london-experience">Planning Your London Experience</h2>
<h3 id="heading-where-to-stay">Where to Stay</h3>
<p>London offers a range of accommodation options to suit every budget:</p>
<ul>
<li><p><strong>Hotels:</strong> Browse options on <a target="_blank" href="https://www.booking.com/">Booking.com</a>, <a target="_blank" href="https://www.agoda.com/">Agoda</a>, or <a target="_blank" href="https://www.airbnb.co.uk/">Airbnb</a> for a comfortable stay near the event venue.</p>
</li>
<li><p><strong>Short-term Rentals:</strong> Consider serviced apartments if you prefer a homier experience during your stay.</p>
</li>
</ul>
<h3 id="heading-getting-there">Getting There</h3>
<p>London is well connected by international air travel:</p>
<ul>
<li><p><strong>Heathrow Airport (LHR)</strong> – The largest and busiest airport, with easy public transit into central London.</p>
</li>
<li><p><strong>Gatwick Airport (LGW)</strong> – A convenient alternative for many international travellers.</p>
</li>
</ul>
<p>Plan your journey ahead to make the most of your time in this vibrant city!</p>
<hr />
<h2 id="heading-visiting-london">Visiting London</h2>
<h3 id="heading-culinary-delights">Culinary Delights</h3>
<p>London’s food scene is as diverse as it is delicious. Whether you’re looking for Michelin-starred restaurants or quirky food markets, here are some recommendations:</p>
<ul>
<li><p><strong>Restaurants:</strong> Try iconic spots in Soho or trendy eateries in Shoreditch.</p>
</li>
<li><p><strong>After-Hours:</strong> Explore vibrant nightlife in Camden or the West End for live music and cocktails.</p>
</li>
<li><p><strong>Coffee Shops:</strong> Recharge at local favorites like Monmouth Coffee or The Attendant for a caffeine boost.</p>
</li>
</ul>
<h3 id="heading-weekend-plans">Weekend Plans</h3>
<p>When you’re not immersed in conference sessions, take some time to explore London’s rich history and culture:</p>
<p><img src="https://www.londonperfect.com/cdn-cgi/image/format=auto,width=1256/https://www.londonperfect.com/g/photos/upload/sml_342226895-1498585820-london-eye-guide.jpg" alt="London Eye" /></p>
<ul>
<li><p><strong>The British Museum:</strong> Discover art and antiquities from around the world.</p>
</li>
<li><p><strong>Tower of London:</strong> Step back in time with a visit to this historic fortress.</p>
</li>
<li><p><strong>Buckingham Palace &amp; Changing of the Guard:</strong> A must-see for first-time visitors.</p>
</li>
<li><p><strong>London Eye:</strong> Enjoy panoramic views of the city skyline.</p>
</li>
</ul>
<hr />
<h2 id="heading-visit-the-doctor-droid-booth">Visit the Doctor Droid Booth</h2>
<p>Doctor Droid is the intelligent Slack bot that accelerates incident diagnosis by automatically pinpointing the root cause of production issues. Simply tag the bot in your alert messages, and let it do the heavy lifting!</p>
<p>Stop by our booth at KubeCon + CloudNativeCon Europe 2025 to discover:</p>
<ul>
<li><p><strong>Live Demos:</strong> See Doctor Droid in action and learn how it can transform your incident response.</p>
</li>
<li><p><strong>Puzzles &amp; Giveaways:</strong> Test your skills and win exciting Doctor Droid goodies.</p>
</li>
<li><p><strong>$500 Doctor Droid Credits:</strong> Show this blog at our booth and receive $500 in credits to supercharge your troubleshooting capabilities.</p>
</li>
</ul>
<p>For those interested in a one-on-one demo, pre-book a meeting with us <a target="_blank" href="https://calendly.com/siddarthjain/kubecon-2024-demo">here</a>.</p>
<hr />
<p>Get ready to experience the future of cloud native computing in one of the world’s most exciting cities – London awaits at KubeCon + CloudNativeCon Europe 2025!</p>
<hr />
<p><em>Happy conferencing, and see you in London!</em></p>
]]></content:encoded></item><item><title><![CDATA[Tools can't buy you good MTTR.. but these 3 practices can]]></title><description><![CDATA[Context
It’s a scenario we’ve all witnessed: teams equipped with cutting-edge observability tools still struggling to catch issues before customers notice.
They’ve invested heavily in top-tier APM solutions, container and infrastructure monitoring, a...]]></description><link>https://notes.drdroid.io/tools-cant-buy-you-good-mttr-but-these-3-practices-can</link><guid isPermaLink="true">https://notes.drdroid.io/tools-cant-buy-you-good-mttr-but-these-3-practices-can</guid><category><![CDATA[observability]]></category><category><![CDATA[monitoring]]></category><category><![CDATA[incident response]]></category><category><![CDATA[logging]]></category><category><![CDATA[#prometheus]]></category><dc:creator><![CDATA[Siddarth Jain]]></dc:creator><pubDate>Tue, 26 Nov 2024 10:08:58 GMT</pubDate><enclosure url="https://cdn.hashnode.com/res/hashnode/image/upload/v1732615128540/89904969-35c9-4170-9a6e-3b58ce1cbaa0.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h3 id="heading-context">Context</h3>
<p>It’s a scenario we’ve all witnessed: teams equipped with cutting-edge observability tools still struggling to catch issues before customers notice.</p>
<p>They’ve invested heavily in top-tier APM solutions, container and infrastructure monitoring, and log accessibility. Yet, their on-call engineers remain overwhelmed. Incidents happen more frequently than anyone would like, and the spotlight they find themselves in post-incident is never the kind they want.</p>
<p>For engineering teams, being called out for production issues is a tough pill to swallow. The key lies in post-incident action plans that lead to meaningful, systematic improvements.</p>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1732612857673/0f6372d4-a2a8-498b-82cc-e3e71c7709cd.png" alt="Whodunnit - I should know before others" class="image--center mx-auto" /></p>
<p>While production incidents can’t be entirely eliminated, well-thought-out preventive measures can dramatically improve operational health.</p>
<h3 id="heading-tools-are-the-baselinenot-the-answer"><strong>Tools Are the Baseline—Not the Answer</strong></h3>
<p>While tools are essential, they primarily address infrastructure or service-level issues. However, most real-world incidents cascade across multiple stacks, often affecting features, products, or customer experiences—areas that are rarely solved by out-of-the-box tools.</p>
<p>To reduce MTTR, teams need processes that improve detection, diagnosis, and resolution speed.</p>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1732613249893/ffc38b35-e29c-426e-8c57-7f87271c05e1.png" alt="Cascading issues from infrastructure to customer experience" class="image--center mx-auto" /></p>
<h3 id="heading-measures-for-reducing-mttr-drastically">Measures for reducing MTTR drastically:</h3>
<p>Here are three processes that I have seen help teams in improving MTTR significantly:</p>
<ol>
<li><p><strong>Improving Actionability of Alerts (Faster Detection)</strong></p>
<ul>
<li><p>Trustworthy alerts are a cornerstone of effective incident management. Engineers need a single source of truth to detect issues early—before customers or business stakeholders notice.</p>
</li>
<li><p>Poorly configured alerts can destroy this trust, leading teams to rely on escalations from support or business teams instead. Monitoring alert quality is critical. For example, many companies using <a target="_blank" href="https://drdroid.io/doctor-droid-slack-integration">Doctor Droid</a> track alert quality to ensure non-actionable alerts don’t erode confidence in their systems.</p>
</li>
</ul>
</li>
<li><p><strong>Instrumenting Custom Metrics</strong></p>
<ul>
<li><p>Custom metrics are invaluable for tracking operational health and catching issues tied to features and product breakages. Unlike generic service-level metrics, custom metrics provide leading indicators that can help teams spot potential failures before they escalate.</p>
</li>
<li><p>By focusing on metrics relevant to their features and customer experience, teams can gain clarity and react faster.</p>
</li>
</ul>
</li>
<li><p><strong>Faster Fixing Through Runbooks and Quick Links</strong></p>
<ul>
<li><p>Developer experience during on-call is often overlooked. Simple resources like runbooks or quick links for known issues can dramatically reduce the cognitive load on engineers.</p>
</li>
<li><p>For example, a link to a pre-built log query can save critical minutes during an incident. These tools empower teams to pinpoint issues faster, enabling quicker resolutions.</p>
</li>
</ul>
</li>
</ol>
<h3 id="heading-conclusion">Conclusion:</h3>
<p>No matter how much you spend on tools, improving MTTR requires engineering investment in processes that enhance detection, diagnosis, and resolution. Custom metrics, actionable alerts, and developer-friendly resources are what truly make the difference.</p>
<p>Engineering teams that focus on these practices find themselves more prepared, more resilient, and better positioned to handle the inevitable challenges of production.</p>
<p><strong>Want to monitor your alerting quality and improve MTTR? Doctor Droid has helped 40+ companies take their incident management to the next level. Get started for free and improve your alerts today!</strong></p>
]]></content:encoded></item><item><title><![CDATA[KubeCon + CloudNativeCon India 2024 Guide -- Delhi]]></title><description><![CDATA[Doctor Droid’s Guide to KubeCon + CloudNativeCon India 2024
Welcome to our complete guide for navigating KubeCon + CloudNativeCon India 2024! Here’s all you need to know to make the most of this event in Delhi, India, from December 11-12, 2024. Don’t...]]></description><link>https://notes.drdroid.io/kubecon-cloudnativecon-india-2024-guide-delhi</link><guid isPermaLink="true">https://notes.drdroid.io/kubecon-cloudnativecon-india-2024-guide-delhi</guid><category><![CDATA[kubeconIN]]></category><category><![CDATA[Kubecon]]></category><category><![CDATA[india]]></category><dc:creator><![CDATA[Jayesh Sadhwani]]></dc:creator><pubDate>Tue, 12 Nov 2024 19:42:07 GMT</pubDate><content:encoded><![CDATA[<h2 id="heading-doctor-droids-guide-to-kubecon-cloudnativecon-india-2024"><strong>Doctor Droid’s Guide to KubeCon + CloudNativeCon India 2024</strong></h2>
<p>Welcome to our complete guide for navigating KubeCon + CloudNativeCon India 2024! Here’s all you need to know to make the most of this event in Delhi, India, from December 11-12, 2024. Don’t forget to visit Doctor Droid at Booth—show us this blog for a chance to receive $500 worth of Doctor Droid credits!</p>
<hr />
<h3 id="heading-overview"><strong>Overview</strong></h3>
<p>KubeCon + CloudNativeCon is the leading conference for Kubernetes, cloud-native technologies, and open-source solutions. Hosted by the Cloud Native Computing Foundation (CNCF), this event gathers thousands of developers, engineers, and business leaders to exchange knowledge, network, and discover the future of cloud-native ecosystems.</p>
<hr />
<h3 id="heading-access-types"><strong>Access Types</strong></h3>
<p><strong>KubeCon + CloudNativeCon India 2024</strong> offers one access pass to cater to all types of attendees:</p>
<p>Includes access to all sessions, keynotes, and co-located events</p>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1731366135812/1aac2ac0-ed1f-493b-a6a7-156a99db3252.png" alt class="image--center mx-auto" /></p>
<ul>
<li><p><strong>Academic Pass</strong>: Discounted pass for students looking to dive into cloud-native technologies.</p>
</li>
<li><p><strong>Individual pass</strong>: Perfect for attendees who are paying for the conference by themselves.</p>
</li>
</ul>
<p>Be sure to review the registration options on the official <strong>KubeCon + CloudNativeCon India 2024</strong> <a target="_blank" href="https://events.linuxfoundation.org/kubecon-cloudnativecon-india"><strong>website</strong></a>.</p>
<h3 id="heading-exclusive-kubecon-cloudnativecon-india-2024-discount-save-20-on-tickets-courtesy-of-doctor-droid"><strong>Exclusive KubeCon + CloudNativeCon India 2024 Discount – Save 20% on Tickets, Courtesy of Doctor Droid!</strong></h3>
<p>That’s right—Doctor Droid is a proud sponsor of KubeCon + CloudNativeCon India 2024, and we’re hooking you up with an exclusive 20% discount on your tickets! Here’s how to secure your spot and save big:</p>
<ol>
<li><p><strong>Fill out this</strong> <a target="_blank" href="https://forms.gle/4obXDnZ41nQBNMH18"><strong>google form</strong></a> with your basic info.</p>
</li>
<li><p><strong>Get your 20% discount code</strong> delivered straight to your inbox.</p>
</li>
<li><p><strong>Register for KubeCon + CloudNativeCon India 2024</strong> and enjoy the savings!</p>
</li>
</ol>
<p>Don’t sit on this—grab your discount before it’s gone!</p>
<h3 id="heading-speakers-amp-tracks-at-kubecon-cloudnativecon-india-2024"><strong>Speakers &amp; Tracks at KubeCon + CloudNativeCon India 2024</strong></h3>
<p><strong>KubeCon + CloudNativeCon</strong> features keynotes from industry leaders and breakout sessions across multiple tracks, including:</p>
<ul>
<li><p><strong>Kubernetes Operations</strong>: Topics around the deployment, scaling, and security of Kubernetes.</p>
</li>
<li><p><strong>AI and Machine Learning</strong>: How AI and ML are transforming Kubernetes.</p>
</li>
<li><p><strong>Edge Computing</strong>: Use cases and solutions for Kubernetes at the edge.</p>
</li>
<li><p><strong>Cloud Security</strong>: Tools and practices to ensure security in a cloud-native environment.</p>
</li>
</ul>
<h3 id="heading-planning-for-kubecon-cloudnativecon-india-2024-delhi-logistics"><strong>Planning for KubeCon + CloudNativeCon India 2024 Delhi Logistics</strong></h3>
<h4 id="heading-stays-near-delhi">Stays near Delhi</h4>
<p>Delhi has several convenient options for accommodation. Whether you prefer hotels or Airbnbs, here are a few recommendations:</p>
<ul>
<li><p><strong>Hotels</strong>: You can find a number hotels on <a target="_blank" href="http://booking.com"><strong>booking.com</strong></a>, <a target="_blank" href="https://agoda.com">agoda.com</a> or <a target="_blank" href="https://makemytrip.com">makemytrip.com</a></p>
</li>
<li><p><strong>Airbnb</strong>: Airbnb has a healthy number of properties available in the city</p>
</li>
</ul>
<h4 id="heading-airports-nearby">Airports nearby</h4>
<ul>
<li><strong>Indira Gandhi International Airport (DEL)</strong> is the nearest airport from the convention center</li>
</ul>
<h3 id="heading-visiting-delhi"><strong>Visiting Delhi</strong></h3>
<h4 id="heading-restaurants">Restaurants</h4>
<p>Delhi offers a diverse culinary scene. Recommended spots include:</p>
<h4 id="heading-after-hour-locations">After Hour Locations</h4>
<h4 id="heading-coffee-shops">Coffee Shops</h4>
<p>Need a caffeine boost? Check out:</p>
<ul>
<li><p><strong>Blue Tokai Coffee</strong>: Known for freshly roasted coffee.</p>
</li>
<li><p><strong>Third Wave Coffee</strong>: Multiple locations to elevate your coffee experience</p>
</li>
</ul>
<h4 id="heading-weekend-plans">Weekend Plans</h4>
<p>Consider exploring Delhi’s culture and history over the weekend. Popular options include:</p>
<ul>
<li><p><strong>Red Fort</strong>: A red coloured fort in the old Delhi to experience the culture and food of Delhi</p>
<p>  <img src="https://lh3.googleusercontent.com/p/AF1QipMzixUK6xvfX9g6zKxOepzWuvo1AfY43mJZAC9g=s1360-w1360-h1020-rw" alt="Photo of Red Fort Lahori Gate" /></p>
</li>
<li><p><strong>Akshardham Temple</strong>: A stunning modern temple complex showcasing India’s rich cultural heritage through intricate architecture, gardens, and a beautiful water show. It’s a must-visit for its grandeur and serenity.</p>
</li>
</ul>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1731440167898/44a32c59-d444-4824-9ad8-c0d35de566e0.png" alt class="image--center mx-auto" /></p>
<p>There are many more places to explore and you can checkout while in delhi</p>
<ol>
<li><p><strong>India Gate</strong>: A monumental war memorial built in honor of Indian soldiers, surrounded by beautiful lawns, making it a great spot for a peaceful evening stroll.</p>
</li>
<li><p><strong>Lotus Temple</strong>: Known for its stunning flower-like shape, this temple is a peaceful place for meditation.</p>
</li>
<li><p><strong>National Museum</strong>: A treasure trove of India’s art, history, and culture spanning millennia.</p>
</li>
<li><p><strong>Dilli Haat</strong>: A market with traditional handicrafts and foods from various Indian states. A must-visit for authentic Indian souvenirs.</p>
</li>
</ol>
<h3 id="heading-visit-doctor-droid-booth"><strong>Visit Doctor Droid Booth</strong></h3>
<p>Doctor Droid is a root cause identification slack bot which can assist on-call engineers diagnose incidents and find root cause really fast. All you need to do is reply to your alert message in slack by tagging the bot. If you are interested to get a demo and explore more about Doctor Droid, visit us at Booth in the venue!</p>
<p><a target="_blank" href="https://calendly.com/siddarthjain/kubecon-2024-demo"><strong>Pre-book a meeting with us using this link.</strong></a></p>
<p>Stop by Booth to discover how Doctor Droid’s automated RCA can help you debug &amp; fix your production issues faster! What else is up for grabs at the event?</p>
<ul>
<li><p><strong>Puzzles &amp; Goodies</strong>: Test your mental muscle and win some amazing gifts.</p>
</li>
<li><p><strong>$500 Doctor Droid Credits</strong>: Show this blog at our booth to receive $500 in Doctor Droid credits!</p>
</li>
</ul>
]]></content:encoded></item><item><title><![CDATA[KubeCon + CloudNativeCon North America 2024 Guide -- Salt Lake City, Utah]]></title><description><![CDATA[Doctor Droid’s Guide to KubeCon + CloudNativeCon North America 2024
Welcome to our complete guide for navigating KubeCon + CloudNativeCon North America 2024! Here’s all you need to know to make the most of this event in Salt Lake City, Utah, from Nov...]]></description><link>https://notes.drdroid.io/kubecon-cloudnativecon-north-america-2024-guide-salt-lake-city-utah</link><guid isPermaLink="true">https://notes.drdroid.io/kubecon-cloudnativecon-north-america-2024-guide-salt-lake-city-utah</guid><category><![CDATA[Kubecon]]></category><category><![CDATA[#cloudnativecon]]></category><dc:creator><![CDATA[Siddarth Jain]]></dc:creator><pubDate>Mon, 04 Nov 2024 19:01:00 GMT</pubDate><enclosure url="https://cdn.hashnode.com/res/hashnode/image/upload/v1730746786664/8e2af6e0-4460-4e9c-983a-63cc29ee2dc7.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h2 id="heading-doctor-droids-guide-to-kubecon-cloudnativecon-north-america-2024"><strong>Doctor Droid’s Guide to KubeCon + CloudNativeCon North America 2024</strong></h2>
<p>Welcome to our complete guide for navigating KubeCon + CloudNativeCon North America 2024! Here’s all you need to know to make the most of this event in Salt Lake City, Utah, from November 12-15, 2024. Don’t forget to visit Doctor Droid at Booth Q45—show us this blog for a chance to receive $500 worth of Doctor Droid credits!</p>
<hr />
<h3 id="heading-overview">Overview</h3>
<p>KubeCon + CloudNativeCon is the leading conference for Kubernetes, cloud-native technologies, and open-source solutions. Hosted by the Cloud Native Computing Foundation (CNCF), this event gathers thousands of developers, engineers, and business leaders to exchange knowledge, network, and discover the future of cloud-native ecosystems.</p>
<hr />
<h3 id="heading-access-types">Access Types</h3>
<p><strong>KubeCon + CloudNativeCon North America 2024</strong> offers different access passes to cater to all types of attendees:</p>
<ul>
<li><p><strong>Full Access Pass</strong>: Includes access to all sessions, keynotes, and co-located events</p>
<p>  <img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1730744336800/322260fe-30a6-4f60-90df-287510cd7a30.png" alt class="image--center mx-auto" /></p>
</li>
<li><p><strong>KubeCon + CloudNativeCon Only Pass:</strong> Includes access to all sessions, keynotes, excluding co-located events</p>
<p>  <img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1730744324464/55a245ea-b06f-462a-859c-a756f318459c.png" alt class="image--center mx-auto" /></p>
</li>
<li><p><strong>Academic Pass</strong>: Discounted pass for students looking to dive into cloud-native technologies.</p>
</li>
<li><p><strong>Individual pass</strong>: Perfect for attendees who are paying for the conference by themselves.</p>
</li>
</ul>
<p>Be sure to review the registration options on the official <strong>KubeCon + CloudNativeCon North America 2024</strong> <a target="_blank" href="https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/">website</a>.</p>
<h3 id="heading-exclusive-kubecon-cloudnativecon-north-america-2024-discount-save-20-on-tickets-courtesy-of-doctor-droid"><strong>Exclusive KubeCon + CloudNativeCon North America 2024 Discount – Save 20% on Tickets, Courtesy of Doctor Droid!</strong></h3>
<p>That’s right—Doctor Droid is a proud sponsor of KubeCon + CloudNativeCon North America 2024, and we’re hooking you up with an exclusive 20% discount on your tickets! Here’s how to secure your spot and save big:</p>
<ol>
<li><p><strong>Fill out this</strong> <a target="_blank" href="https://forms.gle/z3VSYER6RH97Ruv27"><strong>google form</strong></a> with your basic info.</p>
</li>
<li><p><strong>Get your 20% discount code</strong> delivered straight to your inbox.</p>
</li>
<li><p><strong>Register for KubeCon + CloudNativeCon North America 2024</strong> and enjoy the savings!</p>
</li>
</ol>
<p>Don’t sit on this—grab your discount before it’s gone!</p>
<hr />
<h3 id="heading-co-located-events-at-kubecon-cloudnativecon-north-america-2024"><strong>Co-Located Events at KubeCon + CloudNativeCon North America 2024</strong></h3>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1730743508197/651c827a-8bdd-4102-84a0-81d149df7b70.png" alt class="image--center mx-auto" /></p>
<p>Expand your KubeCon + CloudNativeCon experience by joining co-located events, each tailored to specific interests within the cloud-native realm:</p>
<ul>
<li><p><a target="_blank" href="https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/appdevelopercon/"><strong>AppDeveloperCon</strong></a>: Focuses on tools and techniques for building cloud-native applications.</p>
</li>
<li><p><a target="_blank" href="https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/argocon/"><strong>ArgoCon</strong></a>: Dive into Argo workflows, events, and continuous delivery for Kubernetes.</p>
</li>
<li><p><a target="_blank" href="https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/backstagecon/"><strong>BackstageCon</strong></a>: Explore Backstage’s developer portal and best practices for engineering platforms.</p>
</li>
<li><p><a target="_blank" href="https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cilium-ebpf-day/"><strong>Cilium + eBPF Day</strong></a>: A deep dive into Cilium and eBPF for networking and security.</p>
</li>
<li><p><a target="_blank" href="https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cloud-native-kubernetes-ai-day/"><strong>Cloud Native &amp; Kubernetes AI Day</strong></a>: Discuss AI and ML workloads on Kubernetes.</p>
</li>
<li><p><a target="_blank" href="https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cloud-native-startupfest/"><strong>CloudNative StartupFest</strong></a>: Networking and insights for startup founders and innovators.</p>
</li>
<li><p><a target="_blank" href="https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cloud-native-university/"><strong>Cloud Native University</strong></a>: Educational sessions on the fundamentals of cloud-native technologies.</p>
</li>
<li><p><a target="_blank" href="https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/data-on-kubernetes-day/"><strong>Data on Kubernetes Day</strong></a>: Explore data management practices and tools for Kubernetes.</p>
</li>
<li><p><a target="_blank" href="https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/envoycon/"><strong>EnvoyCon</strong></a>: Dedicated to the Envoy proxy community, focusing on networking and observability.</p>
</li>
<li><p><a target="_blank" href="https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/istio-day/"><strong>Istio Day</strong></a>: Learn about Istio and service mesh technologies in cloud-native environments.</p>
</li>
<li><p><a target="_blank" href="https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/kubernetes-on-edge-day/"><strong>Kubernetes on Edge Day</strong></a>: Explore the role of Kubernetes in edge computing.</p>
</li>
<li><p><a target="_blank" href="https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/observability-day/"><strong>Observability Day</strong></a>: A day centered around observability tools and practices in the cloud.</p>
</li>
<li><p><a target="_blank" href="https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/openfeature-summit/"><strong>OpenFeature Summit</strong></a>: Discussions on feature flagging and experimentation in cloud-native setups.</p>
</li>
<li><p><a target="_blank" href="https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/opentofu-day/"><strong>OpenTofu Day</strong>:</a> Open-source infrastructure management and IaC best practices.</p>
</li>
<li><p><a target="_blank" href="https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/platform-engineering-day/"><strong>Platform Engineering Day</strong></a>: Dedicated to platform engineering in cloud-native environments.</p>
</li>
<li><p><a target="_blank" href="https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/wasmcon/"><strong>WasmCon</strong></a>: Focuses on WebAssembly and its role in cloud-native development.</p>
</li>
</ul>
<hr />
<h3 id="heading-kubecon-cloudnativecon-north-america-2024-unofficial-conference-parties">KubeCon + CloudNativeCon North America 2024 Unofficial Conference Parties</h3>
<p>After a day of learning and networking, unwind at the <strong>official</strong> conference parties.</p>
<ul>
<li><p>Check out <a target="_blank" href="https://conferenceparties.com/kubecon24/">Conference Parties</a> for the latest info on social events.</p>
</li>
<li><p>Check out <a target="_blank" href="https://lu.ma/salt-lake-city">events on lu.ma</a> too to keep exploring.</p>
</li>
</ul>
<h3 id="heading-speakers-amp-tracks-at-kubecon-cloudnativecon-north-america-2024">Speakers &amp; Tracks at KubeCon + CloudNativeCon North America 2024</h3>
<p><strong>KubeCon + CloudNativeCon</strong> features keynotes from industry leaders and breakout sessions across multiple tracks, including:</p>
<ul>
<li><p><strong>Kubernetes Operations</strong>: Topics around the deployment, scaling, and security of Kubernetes.</p>
</li>
<li><p><strong>AI and Machine Learning</strong>: How AI and ML are transforming Kubernetes.</p>
</li>
<li><p><strong>Edge Computing</strong>: Use cases and solutions for Kubernetes at the edge.</p>
</li>
<li><p><strong>Cloud Security</strong>: Tools and practices to ensure security in a cloud-native environment.</p>
</li>
</ul>
<hr />
<h3 id="heading-planning-for-kubecon-cloudnativecon-north-america-2024-utah-logistics">Planning for KubeCon + CloudNativeCon North America 2024 Utah Logistics</h3>
<h4 id="heading-stays-near-salt-lake-city">Stays near Salt Lake City</h4>
<p>Salt Lake City has several convenient options for accommodation. Whether you prefer hotels or Airbnbs, here are a few recommendations:</p>
<ul>
<li><p><strong>Hotels</strong>: As of now, all hotels on booking.com as well as most other websites are sold out right now.</p>
</li>
<li><p><strong>Airbnb</strong>: Airbnb still has a healthy number of properties available in the city, although most of them are not within walking distance of the Convention Center.</p>
</li>
<li><p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1730745479198/44962b15-f040-431b-b3d6-5439b2463c6b.png" alt class="image--center mx-auto" /></p>
</li>
</ul>
<h4 id="heading-airports-nearby">Airports nearby</h4>
<ul>
<li><strong>Salt Lake City International Airport (SLC)</strong> is the nearest airport, just 5 miles from downtown.</li>
</ul>
<hr />
<h3 id="heading-visiting-salt-lake-city">Visiting Salt Lake City</h3>
<h4 id="heading-restaurants">Restaurants</h4>
<p>Salt Lake City offers a diverse culinary scene. Recommended spots include:</p>
<ul>
<li><p><strong>Red Iguana</strong>: Known for authentic Mexican cuisine.</p>
</li>
<li><p><strong>The Copper Onion</strong>: A top choice for American fare.</p>
</li>
<li><p><strong>Takashi</strong>: Excellent sushi in the heart of the city.</p>
</li>
</ul>
<h4 id="heading-after-hour-locations">After Hour Locations</h4>
<ul>
<li><p><strong>Beer Bar</strong>: Perfect for a laid-back evening with a variety of beers.</p>
</li>
<li><p><strong>The Bayou</strong>: Offers an extensive selection of brews and Cajun-style food.</p>
</li>
</ul>
<h4 id="heading-coffee-shops">Coffee Shops</h4>
<p>Need a caffeine boost? Check out:</p>
<ul>
<li><p><strong>La Barba Coffee</strong>: Known for artisanal coffee.</p>
</li>
<li><p><strong>Publik Coffee Roasters</strong>: A great spot to unwind with quality brews.</p>
</li>
</ul>
<h4 id="heading-weekend-plans">Weekend Plans</h4>
<p>Consider exploring Utah’s natural beauty over the weekend. Popular options include:</p>
<ul>
<li><p><strong>Bonneville Salt Flats</strong>: A unique desert landscape.</p>
<p>  <img src="https://images.ctfassets.net/0wjmk6wgfops/17oZGsiEevOkg7tpUeaFG0/67bf9df09bff4cbb4136881fa771b789/AdobeStockSaltFlats.jpeg?w=1200&amp;h=630&amp;f=center&amp;fit=fill" alt="Bonneville Salt Flats | Utah.com" /></p>
</li>
<li><p><strong>Big Cottonwood Canyon</strong>: Ideal for scenic drives and hikes.</p>
<p>  <img src="https://dynamic-media-cdn.tripadvisor.com/media/photo-o/15/4a/aa/ab/big-cottonwood-canyon.jpg?w=1200&amp;h=1200&amp;s=1" alt="BIG COTTONWOOD CANYON: All You Need to Know BEFORE You Go" /></p>
</li>
</ul>
<hr />
<h3 id="heading-visit-doctor-droid-booth-q45">Visit Doctor Droid Booth - Q45</h3>
<p>Doctor Droid is a root cause identification slack bot which can assist on-call engineers diagnose incidents and find root cause really fast. All you need to do is reply to your alert message in slack by tagging the bot. If you are interested to get a demo and explore more about Doctor Droid, visit us at Booth Q45 in the venue!</p>
<p><a target="_blank" href="https://calendly.com/siddarthjain/kubecon-2024-demo">Pre-book a meeting with us using this link.</a></p>
<p>Stop by Booth Q45 to discover how Doctor Droid’s automated RCA can help you debug &amp; fix your production issues faster! What else is up for grabs at the event?</p>
<ul>
<li><strong>Puzzles &amp; Goodies</strong>: Test your mental muscle and win some amazing gifts.</li>
</ul>
<ul>
<li><strong>$500 Doctor Droid Credits</strong>: Show this blog at our booth to receive $500 in Doctor Droid credits!</li>
</ul>
<hr />
]]></content:encoded></item></channel></rss>