<?xml version="1.0" encoding="UTF-8"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>Ritwik Reddy · Notes &amp; reflections</title><link>https://ritwikreddy.com/blog</link><description>Engineering notes on infrastructure, automation, and reliable systems.</description><language>en-us</language><atom:link href="https://ritwikreddy.com/feed.xml" rel="self" type="application/rss+xml"/><item><title>Recursive Self-Learning in AI Agents: What Dream-RSI Changes</title><link>https://ritwikreddy.com/notes/dream-rsi-learning-how-to-search</link><guid isPermaLink="true">https://ritwikreddy.com/notes/dream-rsi-learning-how-to-search</guid><description>An analysis of recursive self-learning in AI agents through Dream-RSI: how replay improves exploration, what the results show, and where self-improvement falls short.</description><pubDate>Fri, 18 Sep 2026 00:00:00 GMT</pubDate><content:encoded><![CDATA[<h2 id="heading-intro-the-experiment-after-the-experiment">Intro: the experiment after the experiment</h2>
<p>An AI agent tries an idea. It fails. It tries another. Eventually, something works.</p>
<p>We celebrate the successful output. But the expensive part is often everything that happened before it: the branches pursued for too long, the promising failures abandoned too early, the experiments repeated because nobody improved the process choosing them.</p>
<p>What if those attempts could teach an agent how to spend its next thousand attempts?</p>
<p>That is the interesting question behind <strong>Dream-RSI: Recursive Self-Improvement through Evolving Worlds</strong>, a September 2026 research paper by Tong Zheng and colleagues. Its central proposal is to turn recorded discovery histories into environments where alternative exploration strategies can be tested cheaply. The system then deploys a better strategy, collects more experience, and repeats. <a href="https://arxiv.org/abs/2609.14858v1">Read the paper, version 1</a>.</p>
<p>The phrase <em>recursive self-improvement</em> invites enormous expectations. Here, its meaning is specific: the system revises the code that decides where to search. The underlying language models remain fixed.</p>
<p>That scope makes the research useful. It identifies a part of agent performance that engineers can isolate, measure, and improve: <strong>how an agent allocates effort when the answer is not yet known.</strong></p>
<h2 id="heading-inside-the-research-learning-how-to-explore">Inside the research: learning how to explore</h2>
<h3 id="heading-what-does-recursive-self-learning-mean-here">What does recursive self-learning mean here?</h3>
<p>In this article, <em>recursive self-learning</em> means a feedback loop where experience improves the process used to gather the next round of experience. The paper’s more precise term is <em>recursive self-improvement</em>: Dream-RSI changes its exploration policy, rather than retraining its underlying model. It is also different from asking a model to critique an answer repeatedly; the revisions are evaluated against recorded experimental outcomes.</p>
<h3 id="heading-a-capable-worker-still-needs-a-good-strategy">A capable worker still needs a good strategy</h3>
<p>Imagine a coding agent optimizing a slow program. It could refine the best implementation, try a different algorithm, investigate a failed experiment, or run several approaches in parallel. It also needs to decide when further work is unlikely to justify its cost.</p>
<p>Those decisions constitute an <em>exploration policy</em>. A strong model can still waste a large budget if that policy repeatedly chooses unproductive work.</p>
<p>Testing a new policy directly is expensive. An individual candidate can be evaluated by running it. A strategy must guide many candidates before its quality becomes apparent. Learning how to conduct experiments can therefore require another expensive layer of experimentation.</p>
<p>Dream-RSI attacks that feedback problem. It separates the agent generating solutions from the controller scheduling its attempts, then uses recorded experience to improve the controller.</p>
<h3 id="heading-history-becomes-something-you-can-run">History becomes something you can run</h3>
<p>During discovery, the system records attempts in a tree. Each attempt preserves its starting relationship to earlier work, the resulting artifact and workspace, evaluation feedback, and score.</p>
<p>A completed tree becomes a replay environment. A candidate policy initially sees only the root. As it chooses to open or extend branches, the simulator reveals the corresponding recorded outcomes. It can test different allocations of effort without asking the coding agent to regenerate those outcomes or rerunning the original evaluations.</p>
<p>For an illustrative example, suppose a recorded search contains three approaches. One improves quickly and then stalls. Another starts poorly but recovers after a repair. A third consumes many attempts without progress. Replay lets different scheduling policies encounter those observations in sequence and be scored on what they choose to pursue.</p>
<p>The policy cannot legitimately peek at future scores. Otherwise, it would be optimizing with hindsight unavailable during real execution. The paper’s replay interface and supplied development prompt restrict decisions to revealed observations.</p>
<p>The distinction from ordinary memory matters. A summary might advise an agent to avoid an approach. Replay lets a controller test whether its decision rules would have allocated effort effectively across the recorded search.</p>
<h3 id="heading-what-actually-improves">What actually improves</h3>
<p>The loop has three stages:</p>
<ol>
<li><strong>Explore:</strong> the current policy schedules real generation and evaluation attempts.</li>
<li><strong>Replay:</strong> completed discovery trees become reusable environments for evaluating alternative policies.</li>
<li><strong>Revise:</strong> a policy-development agent edits the controller code using replay feedback; the best evaluated candidate is selected for the next online round.</li>
</ol>
<p>The discovery model, evaluator, and execution interfaces stay fixed. The controller stays fixed within an online rollout and can change between rollouts. The new rollout expands the history available for the next improvement cycle.</p>
<p>In the main formulation, policy scoring balances the best discovered solution, the number of attempts used, and useful parallel execution. This makes efficiency part of the objective. A policy that reaches a good result by exhausting every branch should not automatically beat one that reaches it with substantially less work.</p>
<p>That design also exposes a human choice: what counts as improvement? The objective’s trade-offs are specified by the system’s designers. Recursive optimization does not remove the need to choose them carefully. <a href="https://arxiv.org/pdf/2609.14858v1">Method: paper sections 2–3</a>.</p>
<h3 id="heading-the-results-worth-paying-attention-to">The results worth paying attention to</h3>
<p>The paper evaluates eight tasks across algorithm engineering, mathematical optimization, and GPU kernel engineering. Its most useful controlled comparison keeps the discovery setup and starting policy the same, while the baseline’s exploration policy remains fixed.</p>
<p>For a Lasso solver, using Gemini-3.1 Pro, Dream-RSI uses <strong>317 discovery-agent calls compared with 550</strong> for fixed exploration. The resulting solver’s average runtime across six held-out datasets falls from <strong>3,587.1 ms to 2,931.0 ms</strong>. That is roughly 42% fewer discovery calls and an 18% reduction in average solver runtime, calculated from the reported values.</p>
<p>There is a revealing detail beneath that average: the Dream-RSI solver is slower on five of those six datasets. Its large improvement on RCV1 drives the overall gain. A better average is meaningful, but it does not establish a uniformly better solution for every workload.</p>
<p>The GPU results show two different benefits. On VGG16, the method reaches comparable performance with 2.43 times fewer generations. On ConvDiv, it achieves 2.09 times the inverse-runtime performance under comparable discovery budgets. One result concerns the cost of finding a solution; the other concerns the quality of the solution found.</p>
<p>The mathematical results are mixed. Dream-RSI improves the sum–difference score and ties the strongest listed circle-packing result. For autocorrelation, where lower is better, its score of 1.456375 is slightly worse than fixed exploration’s 1.456001. Adaptation helps in several settings, but the evidence does not support universal superiority. <a href="https://arxiv.org/pdf/2609.14858v1">Results: Figure 3, Table 1, and Figure 4</a>.</p>
<h3 id="heading-where-the-evidence-stops">Where the evidence stops</h3>
<p>Replay only contains outcomes that were recorded. It cannot establish what an untried branch would have produced. Online generation is stochastic; replay reveals a particular historical result. A strategy that looks strong in that recorded world can still disappoint in a fresh run.</p>
<p>The paper’s selection procedure guarantees that the chosen policy is no worse on the fixed history’s average replay score, because the current policy is among the candidates. <strong>That is a guarantee about replay selection, not a guarantee of continued improvement in the real task.</strong></p>
<p>Cost also needs careful reading. The experiments count discovery-agent calls. Those counts do not by themselves establish equivalent savings in total dollars, tokens, or elapsed time, including policy development. The headline comparison of up to 162 times fewer calls against SimpleTES also involves different model backbones. The same-model fixed-policy comparison gives a cleaner view of the controller’s contribution.</p>
<p>My interpretation is that Dream-RSI demonstrates a promising method for improving exploration in these evaluated discovery settings. It does not demonstrate unlimited self-improvement, general workplace autonomy, or an agent that can rewrite every part of itself successfully.</p>
<h2 id="heading-what-this-could-change-for-ai-agents">What this could change for AI agents</h2>
<p>The possibilities below are extrapolations from the research, rather than outcomes demonstrated by its experiments.</p>
<h3 id="heading-agent-infrastructure-becomes-a-place-to-learn">Agent infrastructure becomes a place to learn</h3>
<p>Agent builders routinely choose retry limits, branching rules, concurrency, and stopping conditions. Dream-RSI suggests that some of those choices could become policies improved through evidence from previous runs.</p>
<p>For a future code-optimization service, the valuable question might be whether a modest gain deserves another refinement, whether a failed implementation deserves repair, or whether an entirely different approach deserves the remaining budget.</p>
<p>If replay transfers reliably to fresh workloads, improvements could come from better orchestration even when the underlying model stays the same. That would give teams another way to improve performance between model upgrades.</p>
<h3 id="heading-operational-history-gains-a-second-purpose">Operational history gains a second purpose</h3>
<p>Logs help explain what happened. Structured discovery traces can additionally support experiments about how work should have been scheduled.</p>
<p>That possibility raises the value of recording relationships between attempts, exact artifact versions, evaluation results, and failures. A final answer and a transcript may be insufficient to reconstruct a useful replay environment.</p>
<p>It also connects to a concern I raised in <a href="https://ritwikreddy.com/notes/when-ai-writes-and-reviews-the-code">When AI Writes the Code, Who Understands It?</a>: faster automated work needs a durable explanation. A system changing its own controller should preserve which version acted, what evidence justified the revision, and whether fresh evaluations confirmed the expected benefit.</p>
<h3 id="heading-better-evaluation-becomes-more-valuable">Better evaluation becomes more valuable</h3>
<p>The evaluated tasks provide comparatively clear feedback: mathematical objectives, numerical checks, and execution performance. Many production tasks have incomplete success criteria and delayed consequences.</p>
<p>A support agent might resolve a ticket quickly while misunderstanding the customer. An infrastructure agent might reduce deployment time while making recovery harder. Optimizing exploration cannot rescue an objective that rewards the wrong outcome.</p>
<p>I would therefore start applying this idea in bounded environments with reproducible evaluations. Candidate controllers should face fresh cases outside the replay history, and their total operating cost should be measured. Deployment would need versioning and rollback so an apparently better controller does not become an irreversible decision.</p>
<h3 id="heading-a-growing-history-can-also-preserve-a-blind-spot">A growing history can also preserve a blind spot</h3>
<p>The most interesting unresolved question is whether the replay collection becomes more representative as it grows.</p>
<p>A policy determines which experiences get collected. Those experiences shape its successor. That creates a useful feedback loop, but also a possible bias: repeatedly searching familiar territory can produce a rich history of a narrow world.</p>
<p>Testing on unfamiliar tasks and preserving opportunities for fresh exploration would help distinguish learning a useful strategy from becoming increasingly skilled at navigating yesterday’s experiments.</p>
<p>Dream-RSI makes a concrete contribution to the agent landscape: it shows how expensive discovery experience can be reused to improve the decisions governing future discovery. The opportunity is substantial wherever experiments are costly and outcomes can be evaluated reliably.</p>
<p>The question I would carry into the next generation of agent systems is this: <strong>after a thousand attempts, has the agent only produced a better answer, or has it learned to spend the next thousand more wisely?</strong></p>
<p><em>Research basis: Tong Zheng et al., Dream-RSI: Recursive Self-Improvement through Evolving Worlds, arXiv:2609.14858v1, submitted September 14, 2026. This article analyzes the supplied version; proposed production applications are the author’s interpretation.</em></p>
]]></content:encoded></item>
<item><title>When AI Writes the Code, Who Understands It?</title><link>https://ritwikreddy.com/notes/when-ai-writes-and-reviews-the-code</link><guid isPermaLink="true">https://ritwikreddy.com/notes/when-ai-writes-and-reviews-the-code</guid><description>AI can write and review a change while human understanding falls behind. A documentation agent could help preserve the knowledge teams need to operate their systems.</description><pubDate>Sat, 12 Sep 2026 00:00:00 GMT</pubDate><content:encoded><![CDATA[<p>A developer describes a feature. An AI agent implements it, another reviews the changes, and automated checks report success. A person scans the summary and approves the merge. The work moves forward with very little human intervention.</p>
<p>There is real value in that workflow. Repetitive implementation and first-pass review can consume attention that engineers could spend elsewhere. But there is a question the green checks cannot answer: who now understands the behavior entering production?</p>
<p>A company can own its repositories and still struggle to explain its systems. My concern is that automating both creation and review can widen that gap when teams remove the moments in which people build understanding. The risk is gradual: changes accumulate faster than anyone develops a reliable mental model of them.</p>
<h2 id="heading-agreement-is-not-understanding">Agreement is not understanding</h2>
<p>AI review can catch mistakes. It can question an implementation, identify missing tests, and offer useful alternatives. GitHub describes Copilot code review as supplementing human review and emphasizes human oversight and validation across its agentic tools. <a href="https://docs.github.com/en/copilot/responsible-use/agents">GitHub’s responsible-use guidance</a></p>
<p>The problem starts when an automated review becomes sufficient evidence for a decision it was never equipped to make. A reviewer may inspect the changed code without knowing that a seemingly redundant check protects an old integration, or that a retry changes the business meaning of an operation.</p>
<p>Using a different model may offer another perspective. It does not supply requirements that nobody documented. Two agents can work from the same incomplete context and produce mutually consistent answers. Their agreement is useful evidence, but it cannot establish that the team has explained the right problem.</p>
<p>Human reviewers have blind spots too. The aim is to preserve someone accountable for connecting implementation choices to the wider system.</p>
<h2 id="heading-the-understanding-gap-becomes-an-operational-problem">The understanding gap becomes an operational problem</h2>
<p>Imagine a hypothetical order service where an agent adds retries after a supplier API times out. The tests pass. An automated review checks exception handling and sees a bounded retry loop. The change appears reasonable.</p>
<p>Later, some customers receive duplicate orders. The supplier sometimes accepted the original request even though its response never reached the service. Retrying created another order. A test built around the assumption that timeout means failure would not expose that distinction.</p>
<p>Now the team needs to answer practical questions. Where is the retry policy defined? Does the supplier support deduplication? Which version introduced the behavior? Can retries be disabled without losing legitimate orders?</p>
<p>Those questions require knowledge across code, external contracts, deployment configuration, and business expectations. A repository search can locate a function. It cannot, by itself, explain a decision whose rationale was never recorded.</p>
<p>The same gap appears outside incidents. A new engineer needs to trace a request through unfamiliar services. Someone fixing a bug needs to know which assumptions must survive the fix. If every explanation must be reconstructed from scratch, faster delivery can leave a growing maintenance burden behind it.</p>
<h2 id="heading-invest-in-an-agent-that-maintains-the-explanation">Invest in an agent that maintains the explanation</h2>
<p>One practical response is to invest in a documentation agent whose job is to keep the system understandable as it changes. Its output should be a maintained knowledge base with evidence, rather than a pile of generated descriptions.</p>
<p>I would begin with a small set of questions every service should answer:</p>
<ul>
<li>What does this service do, and who owns it?</li>
<li>How do requests and data move through it?</li>
<li>Which dependencies and failure modes matter?</li>
<li>How is it deployed, observed, and recovered?</li>
<li>Which design decisions are intentional, and why?</li>
</ul>
<p>The agent could inspect approved repositories, infrastructure definitions, API contracts, tests, and existing decision records. From those sources, it could propose service summaries, dependency maps, runbooks, and onboarding guides.</p>
<p>Each important claim should link to its supporting source and identify the relevant version. “Retries are limited to three attempts” should point to the configuration or implementation. “This protects against duplicate orders” needs evidence of that behavior, not a confident interpretation of a function name.</p>
<h2 id="heading-make-documentation-part-of-the-change">Make documentation part of the change</h2>
<p>The useful integration point is the pull request: the proposed code change and its explanation can be reviewed together. The agent examines what changed, identifies affected documentation, and opens a documentation update alongside the implementation.</p>
<p>For the retry example, it should flag the operational contract: whether requests can be repeated safely, what happens after an ambiguous response, and how an operator can stop further attempts. If it cannot establish those facts, it should ask an explicit question.</p>
<p>After merge, documentation should identify the code revision it describes. Deployment records must then show whether that revision is actually running. The latest repository state and the live system are not necessarily the same thing, particularly during staged rollouts or emergency changes.</p>
<p>Periodic checks can flag stale links, missing owners, changed interfaces, and runbooks affected by new configuration. These checks create review work; they do not prove that every explanation is correct. Teams should record when critical instructions were last exercised, not merely when their wording was regenerated.</p>
<h2 id="heading-do-not-automate-the-same-blind-spot-three-times">Do not automate the same blind spot three times</h2>
<p>There is an obvious objection: if AI writes, reviews, and documents the code, have we simply added another layer of plausible output?</p>
<p>We have, unless the documentation process introduces evidence and human learning. A generated explanation can accurately describe an incorrect implementation. It can also invent a sensible-sounding reason for a decision nobody actually made.</p>
<p>The agent therefore needs to distinguish observed behavior, human-recorded intent, and unresolved inference. An architectural decision record should capture the responsible engineer’s rationale and rejected alternatives. The agent can organize that account, but it should not manufacture the history.</p>
<p>Its access should also match its purpose. Reading relevant code and proposing documentation changes does not require authority to modify production. Sensitive logs, credentials, and customer information should not be copied into broadly accessible guides. The knowledge base should preserve the access boundaries of its sources.</p>
<p>People still need to review consequential changes and explain their effects. Documentation gives them a better starting point; it does not transfer responsibility to the tool that wrote it.</p>
<h2 id="heading-test-whether-people-can-use-it">Test whether people can use it</h2>
<p>During root cause analysis, a documentation agent could help assemble a timeline, locate relevant changes, and connect symptoms to documented dependencies. It should separate observations from hypotheses. A convincing causal story is not the same as a cause supported by evidence.</p>
<p>Google’s SRE guidance treats a postmortem as a record of an incident, its impact, causes, response, and follow-up actions. A documentation agent could help preserve those findings once the team has investigated them. <a href="https://sre.google/sre-book/postmortem-culture/">Postmortem culture</a></p>
<p>The strongest evaluation is practical. Ask an engineer unfamiliar with a service to use its documentation to trace a request, identify its owner, and explain a safe recovery procedure. Run a controlled incident exercise. Check whether the guide helps them find evidence or sends them confidently in the wrong direction.</p>
<p>Measure time to find useful information, incorrect instructions, unanswered questions, and corrections made by maintainers. Page count and generated word count say little about understanding. Start with one service, learn where the agent helps, and expand only when the documentation earns trust in use.</p>
<h2 id="heading-keep-someone-who-can-explain-the-system">Keep someone who can explain the system</h2>
<p>AI-assisted development does not inevitably leave companies unable to understand their code. That outcome depends on how teams organize review, learning, and ownership. Automation can also make explanations easier to produce and maintain.</p>
<p>The investment I would argue for is straightforward: when accelerating code production, fund the work of keeping that code understandable. A documentation agent is one way to reduce that burden, provided its claims remain traceable and people remain involved.</p>
<p>The question after a successful merge should extend beyond whether the checks passed: <strong>can someone on this team explain what changed, why it matters, and what to do when it fails?</strong></p>
]]></content:encoded></item>
<item><title>Giving AI the Authority to Act</title><link>https://ritwikreddy.com/notes/giving-ai-the-authority-to-act</link><guid isPermaLink="true">https://ritwikreddy.com/notes/giving-ai-the-authority-to-act</guid><description>When AI moves from making suggestions to taking action, the question becomes how much responsibility to give it—and where human judgment belongs.</description><pubDate>Thu, 10 Sep 2026 00:00:00 GMT</pubDate><content:encoded><![CDATA[<p>Imagine an online service slowing down during a busy afternoon. An AI assistant examines the symptoms and suggests restarting a component. A person considers the recommendation and decides what to do.</p>
<p>Now give that assistant access to the operational tools. It can investigate, choose a response, and carry it out. The distance between an idea and its consequences becomes much shorter.</p>
<p>That change is exciting. It also makes an old question newly urgent: when we delegate a task, what exactly are we delegating?</p>
<p>The ability to produce a sensible recommendation is one kind of capability. Permission to alter a system that other people depend on is a separate decision. A useful conversation about AI in production needs room for both.</p>
<h2 id="heading-from-answering-to-acting">From answering to acting</h2>
<p>Here, an AI agent means a system that can choose steps, use tools, and respond to what happens as it works toward a goal. The distinction is practical: a predefined workflow follows a route its designers laid out; an agent has some discretion over the route. Anthropic describes this distinction in its engineering guide, which also recommends starting with the simplest approach that meets the task. <a href="https://www.anthropic.com/engineering/building-effective-agents">Building effective agents</a></p>
<p>In production—the systems serving real users—that discretion might involve deciding which logs to examine, preparing a configuration change, or taking an approved operational action. Those activities have very different consequences.</p>
<p>Calling a system “autonomous” tells us surprisingly little. Autonomous over which decisions? With access to whose information? For how long? What happens when the evidence is unclear?</p>
<p>These are questions about the shape of the job we have created.</p>
<h2 id="heading-a-familiar-problem-with-a-different-boundary">A familiar problem, with a different boundary</h2>
<p>Consider a hypothetical agent helping investigate a slow checkout service. It finds that a queue is growing and that a recent deployment changed how work is processed.</p>
<p>Several responses are possible. It could gather more evidence. It could prepare a rollback proposal. It could increase processing capacity. It could disable a feature that appears to be contributing to the delay.</p>
<p>Each choice carries assumptions. More capacity may increase cost without removing the bottleneck. A rollback may conflict with a newer data format. Disabling a feature may improve response times while making checkout unusable for some customers.</p>
<p>The instruction “restore service” does not resolve these tradeoffs. The team still needs to decide which outcomes matter, which constraints must hold, and which uncertainties demand review.</p>
<p>This is why a polished demonstration can leave the central production question unanswered. We have seen what the agent can do under the demonstration’s conditions. We still need to define what it may do when conditions change.</p>
<h2 id="heading-give-actions-a-scope">Give actions a scope</h2>
<p>For this hypothetical service, a starting policy might look like this:</p>
<table>
<thead>
<tr>
<th>Action</th>
<th>Example boundary</th>
</tr>
</thead>
<tbody>
<tr>
<td>Inspect service health</td>
<td>Access only the relevant telemetry, with limits on sensitive data</td>
</tr>
<tr>
<td>Prepare a proposed change</td>
<td>Produce a reviewable change and supporting evidence; do not apply it</td>
</tr>
<tr>
<td>Perform an established recovery action</td>
<td>Use a narrowly authorized procedure with explicit preconditions and limits</td>
</tr>
<tr>
<td>Make an unfamiliar or high-impact change</td>
<td>Pause for a person to assess the proposal</td>
</tr>
</tbody>
</table>
<p>This is an illustrative policy, not a universal classification. Even reading data can be consequential if the data is sensitive or sent somewhere it should not go. A reversible-looking change can still interrupt a customer’s work.</p>
<p>OWASP’s guidance on excessive agency recommends limiting tool functionality and permissions, requiring approval for high-impact actions, and enforcing authorization in downstream systems. The permission check belongs in the mechanism that performs the action; asking the model to remember a boundary is insufficient. <a href="https://genai.owasp.org/llmrisk/llm062025-excessive-agency/">Excessive Agency</a></p>
<p>For our checkout example, the tool might permit a capacity increase only for one service and within a fixed ceiling. An agent’s convincing explanation would not expand that ceiling.</p>
<h2 id="heading-make-human-review-worth-having">Make human review worth having</h2>
<p>A request that says “Approve this fix?” leaves most of the work with the reviewer. They must reconstruct the incident, discover the proposed change, and decide whether the agent has missed anything.</p>
<p>A useful approval request would explain what will change, why it is expected to help, which evidence supports that expectation, what remains uncertain, and how the result will be checked. It would also state what happens if the change makes things worse.</p>
<p>The person should be approving a specific proposal. If that proposal changes materially, the earlier approval should not silently become permission for a different action.</p>
<p>There is a tradeoff here. Review introduces delay and consumes attention. Requiring it for every trivial step can turn oversight into a sequence of habitual clicks. My view is that review is most valuable where consequences or uncertainty justify that attention, with routine actions bounded by controls established in advance.</p>
<p>The design question is therefore quite concrete: what does the person need to see in order to make a meaningful decision?</p>
<h2 id="heading-a-successful-action-still-needs-verification">A successful action still needs verification</h2>
<p>Suppose a tool times out after the agent requests a capacity increase. Did the operation fail, or did it succeed while the response was lost? Repeating the request without checking could create a second change.</p>
<p>This is a familiar distributed-systems problem. AWS describes how idempotent APIs let callers retry operations without causing repeated side effects, using request identifiers and carefully defined behavior. That property is worth considering wherever an agent can retry a consequential tool call. <a href="https://aws.amazon.com/builders-library/making-retries-safe-with-idempotent-APIs/">Making retries safe with idempotent APIs</a></p>
<p>Completion also needs evidence outside the agent’s own account of its work. In our example, a successful tool response might confirm that the requested change was accepted. We would still need to check whether checkout improved and whether another part of the service deteriorated.</p>
<p>Google’s guidance on canary releases describes evaluating a change on a limited portion of a system before expanding it. Applying that principle to agent-driven changes is a design choice: where the system supports it, limit initial exposure and observe the result before proceeding. A small rollout is useful only if the chosen signals can reveal the failures that matter. <a href="https://sre.google/workbook/canarying-releases/">Canarying Releases</a></p>
<h2 id="heading-what-would-justify-more-responsibility">What would justify more responsibility?</h2>
<p>I would want to see more than a collection of successful runs. I would want examples involving missing information, unavailable tools, conflicting signals, and a goal that cannot be achieved within the allowed scope.</p>
<p>Does the agent recognize when it has reached a boundary? Does it leave enough evidence for someone to understand what happened? Can the team interrupt its work and recover from a partial result?</p>
<p>Those questions change what we count as success. Stopping with a precise explanation may be the right outcome. Finishing a task by violating its constraints may be a failure, even if the visible symptom disappears.</p>
<p>The broader lesson reaches beyond AI. Delegation requires us to make expectations explicit. An agent can expose how much of an organization’s operating knowledge still lives in unstated assumptions: who may decide, what counts as acceptable, and when someone must ask for help.</p>
<p>The opportunity is to use that pressure to design better systems of responsibility. Give an agent a useful scope. Make the boundary enforceable. Learn from what happens inside it before changing it.</p>
<p>Then the question becomes more interesting than whether we trust AI in the abstract: <strong>what evidence would make us comfortable giving this system this particular decision?</strong></p>
]]></content:encoded></item>
<item><title>Constraint</title><link>https://ritwikreddy.com/reflections/constraint</link><guid isPermaLink="true">https://ritwikreddy.com/reflections/constraint</guid><description>Good engineering is often the reduction of unnecessary possibility.</description><content:encoded><![CDATA[<p>Constraint is usually misunderstood as limitation. In practice, it is often the thing that makes clarity possible.</p>
<p>A system with infinite options becomes difficult to operate. A life with no structure becomes difficult to direct. Good design removes what does not need to exist.</p>
]]></content:encoded></item>
<item><title>Invisible Systems</title><link>https://ritwikreddy.com/reflections/invisible-systems</link><guid isPermaLink="true">https://ritwikreddy.com/reflections/invisible-systems</guid><description>The best infrastructure disappears into reliability. Most users never notice it exists — only when it fails.</description><content:encoded><![CDATA[<p>The most reliable systems often become invisible. They do not announce themselves. They do not ask for attention. They simply create enough stability for people to focus on the work that actually matters.</p>
<p>This is true in infrastructure, but also in life. The best routines, environments, and habits reduce unnecessary friction until discipline feels less like force and more like design.</p>
<blockquote>
<p>A good system is not the one you admire every day. It is the one you stop worrying about.</p>
</blockquote>
]]></content:encoded></item>
<item><title>Operational Calm</title><link>https://ritwikreddy.com/reflections/operational-calm</link><guid isPermaLink="true">https://ritwikreddy.com/reflections/operational-calm</guid><description>A mature system is not the one with the most features. It is the one that behaves predictably under pressure.</description><content:encoded><![CDATA[<p>Pressure reveals the real architecture of a system. It exposes what was designed carefully and what was only held together by effort.</p>
<p>The same is true for people. Calm is not the absence of intensity. It is the presence of internal structure strong enough to absorb it.</p>
<blockquote>
<p>Maturity is predictability under pressure.</p>
</blockquote>
]]></content:encoded></item>
<item><title>Terraform Drift and the Illusion of Stability</title><link>https://ritwikreddy.com/notes/terraform-drift</link><guid isPermaLink="true">https://ritwikreddy.com/notes/terraform-drift</guid><description>Why infrastructure stops matching intent, how drift enters production systems, and how validation layers reduce surprise.</description><content:encoded><![CDATA[<p>Infrastructure drift is rarely caused by a single catastrophic event. More often, it emerges slowly through manual interventions, emergency fixes, inconsistent provisioning paths, and undocumented operational decisions.</p>
<p>Terraform creates the illusion that infrastructure and intent remain perfectly synchronized. In reality, production systems continuously evolve under operational pressure.</p>
<h2 id="heading-why-drift-becomes-dangerous">Why Drift Becomes Dangerous</h2>
<p>The danger is not merely configuration mismatch. Drift weakens predictability. Once engineers stop trusting deployment state, operational confidence begins collapsing across the system.</p>
<blockquote>
<p>The cheapest failure is the one detected before production behavior diverges.</p>
</blockquote>
<p>Validation layers, governance automation, infrastructure audits, and repeatable deployment workflows become essential once environments scale beyond a few isolated systems.</p>
]]></content:encoded></item></channel></rss>