<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0" xmlns:itunes="http://www.itunes.com/dtds/podcast-1.0.dtd" xmlns:googleplay="http://www.google.com/schemas/play-podcasts/1.0"><channel><title><![CDATA[Compounding Decisions for the Age of AI]]></title><description><![CDATA[Exploring how people, products, and intelligent systems learn, adapt, and compound over time.]]></description><link>https://blog.rememberloop.com</link><image><url>https://blog.rememberloop.com/img/substack.png</url><title>Compounding Decisions for the Age of AI</title><link>https://blog.rememberloop.com</link></image><generator>Substack</generator><lastBuildDate>Wed, 05 Aug 2026 14:28:40 GMT</lastBuildDate><atom:link href="https://blog.rememberloop.com/feed" rel="self" type="application/rss+xml"/><copyright><![CDATA[Sriram Natarajan]]></copyright><language><![CDATA[en]]></language><webMaster><![CDATA[srinatar@substack.com]]></webMaster><itunes:owner><itunes:email><![CDATA[srinatar@substack.com]]></itunes:email><itunes:name><![CDATA[Sriram Natarajan]]></itunes:name></itunes:owner><itunes:author><![CDATA[Sriram Natarajan]]></itunes:author><googleplay:owner><![CDATA[srinatar@substack.com]]></googleplay:owner><googleplay:email><![CDATA[srinatar@substack.com]]></googleplay:email><googleplay:author><![CDATA[Sriram Natarajan]]></googleplay:author><itunes:block><![CDATA[Yes]]></itunes:block><item><title><![CDATA[When AI makes building cheap, the job is knowing what's worth building]]></title><description><![CDATA[A customer discovery call on Monday. A working demo on Wednesday. What changed in between wasn't my speed.]]></description><link>https://blog.rememberloop.com/p/when-ai-makes-building-cheap-the</link><guid isPermaLink="false">https://blog.rememberloop.com/p/when-ai-makes-building-cheap-the</guid><dc:creator><![CDATA[Sriram Natarajan]]></dc:creator><pubDate>Sun, 02 Aug 2026 18:05:47 GMT</pubDate><content:encoded><![CDATA[<p>On a Monday I sat in on a one-hour call with a customer and mostly listened. It was a discovery call, so my job was to watch them work. Three of their people took turns showing me how they solve their day-to-day problems in a different way. By Wednesday I had a 30-minute demo of our product that walked their own workflows back to them.</p><p>The 48 hours is the part people ask about. It wasn&#8217;t sprint energy or an all-nighter. The call handed me almost everything I needed, and I then used AI to automate the pain that users showed me. I didn&#8217;t have to add it to a sprint backlog, run a prioritization debate, or wait for the next planning cycle.</p><h2>What the call actually gave me</h2><p>I stopped thinking of it as &#8220;the customer uses tool x to do y things.&#8221; Three different people used one of our competitor&#8217;s tools, each for their own job, and none of them thought they shared it with the others.</p><p>Their marketing person had the workflow that mattered most. After launching a personalization campaign, she needed a way to confirm that their targeted users saw the personalized offer. That meant pulling a list of users, checking each one against five attributes in a separate system, and watching session replays one at a time to catch the mismatches. Minutes per user, thousands of users, weeks of work per campaign. She got the targeting right in the end, by hand, at a cost that made me wince when she described it.</p><p>Their support person and their product manager had their own routines too, both narrower. But the marketing workflow was the one I kept thinking about on the drive back, because it was weeks of a smart person doing by hand something a machine should do in seconds.</p><p>I wrote each person down as a workflow, not a feature request. A person, and the exact sequence of steps they take.</p><h2>Why workflows instead of notes</h2><p>The old version of me would have left that call with two pages of notes about what the customer wants, and the notes would have flattened all three of those people into one line: they need better analytics. Then I would have taken that line to the next sprint planning meeting and argued with my engineering team about where it ranked against everything else.</p><p>Writing it as a workflow is different because it keeps the steps in order and keeps the persona attached to them. Once I had &#8220;marketing: launch campaign, pull the user list, check five attributes each, watch replays for mismatches,&#8221; I didn&#8217;t have a requirement anymore. I had a script. The demo is that same sequence, run in our product, with AI collapsing the slow middle.</p><h2>The part AI actually changed</h2><p>When I did the demo of the prototype, I walked the marketing workflow back to her out loud. The weeks of manual checking, the five attributes, the replays one at a time. Then I showed that in our product, you can ask in plain English the same question she used to answer by hand: show me the users who saw this offer but didn&#8217;t meet the criteria. A few seconds later, there was the list.</p><p>My job on the call was to see that her weeks of manual effort were the thing worth automating. The build was fast because the hard part was noticing, not coding.</p><p>I put that moment last in the demo on purpose, because it was the one I wanted the room to feel. Before it, I matched the workflows they already trusted, step for step, so they knew we could do what they rely on today. I didn&#8217;t pitch our product as a replacement for the system they treat as their source of truth, because we aren&#8217;t one. You match what they trust first, then you show them the thing that used to be impossible.</p><h2>Why it&#8217;s repeatable</h2><p>The 48 hours worked because almost none of it was invented from scratch. The call gave me the people and their workflows. The workflows became the demo. AI turned the slowest workflow into a single query. If I get another discovery call next month, I run the same steps, and I&#8217;ll be about as fast, because the speed was never mine. It was in the method.</p><p>When AI drastically shrinks your build time, it is important for you to have clarity on what is worth building. As an AI product builder, I constantly find myself asking: does building this truly solve the user&#8217;s problem, and does it capture value? This is what building looks like for me now.</p><p>You sit with a real user and watch where they lose hours repeatedly. You write down those exact steps precisely enough that AI can do them. The best demo is just their own job, with the slow part taken out. </p><p>You need to build the judgement skill to decide what to build and then capture their workflows precisely before you leave the room.</p>]]></content:encoded></item><item><title><![CDATA[How I use AI after a product review with senior executives]]></title><description><![CDATA[I capture the decisions and the options we rejected, and my AI won't let me quietly reopen them.]]></description><link>https://blog.rememberloop.com/p/how-i-use-ai-after-a-product-review</link><guid isPermaLink="false">https://blog.rememberloop.com/p/how-i-use-ai-after-a-product-review</guid><dc:creator><![CDATA[Sriram Natarajan]]></dc:creator><pubDate>Sun, 26 Jul 2026 17:50:24 GMT</pubDate><content:encoded><![CDATA[<p>In March, I sat in a product review with my executive leaders and watched them take apart a product mock for a 0-1 product I had spent weeks on. They were not unkind about it. They just kept pointing at the screen and saying what they actually saw, which was not the value I thought I was delivering.</p><p>They said things like &#8220;the value didn&#8217;t jump out&#8221; and &#8220;I&#8217;m reading metric definitions instead of seeing the point.&#8221; Over about an hour they said a dozen specific things about what was wrong and what they&#8217;d want instead.</p><p>The old version of me would have walked out with a page of notes, written a tidy summary that afternoon, and let it slowly go stale. The summary would say what the meeting was about, but it would not change what we built two months later.</p><p>This time I did something different. I fed every sentence they said into Claude, close to verbatim, and asked it to help me turn the verbatim into a small set of decisions rather than a summary. Each decision was written the same way: here&#8217;s what we decided, here&#8217;s why, and here are the options we rejected and the reasons / principles on why we rejected them.</p><p>Those third and fourth fields, the rejected options and the reasons for rejection, ended up mattering most, though it took me a few weeks to see why.</p><p>The raw input is the executives&#8217; actual words, not my paraphrase of them. In most cases, paraphrases are where the signal leaks out. When I write &#8220;they wanted the overview to be strategic,&#8221; I&#8217;ve already smoothed off the specific thing they said, which was that the highest-value view should lead the page and everything operational should sit under it. The verbatim version keeps the edge. So step one is just capturing what was said, precisely, and handing that to the AI as the source material.</p><p>During that hour-long product review, we collectively made four key product decisions. One of them was about what the product leads with: we decided to lead with the view the executives actually care about, and we rejected framing it as another analytics dashboard, because that puts us right next to the incumbent everyone already benchmarks against. The rejected option is written down with the same weight as the chosen one, and that turned out to matter a lot in my later decisions.</p><p>I&#8217;d initially kept the rejected-option field for the sake of completeness, the kind of thing you write but never read.</p><p>I then set up my Claude so that every future AI session I run on this product preloads this decisions file by default. It&#8217;s the first thing in the project&#8217;s context. So when I open a new session in June and start exploring a specific direction, and that direction happens to be something we already killed in March, the AI stops me right there. It tells me we rejected this option, on this date, for this reason.</p><p>Without the &#8220;alternatives rejected&#8221; field, that doesn&#8217;t happen. The AI only knows what we chose, so it cheerfully re-proposes the thing we already threw out, and I lose an afternoon rediscovering why it was a bad idea the first time. With the field, the closed questions stay closed. The meeting stops being a memory I have to defend in every new conversation and starts being a constraint the system enforces for me.</p><p>I initially built the decisions file as documentation, but now I use it as a way to keep me and the AI from reopening settled questions.</p><p>Six months on, nearly every decision in that product file traces back to that one hour in March. Not because the executives changed my mind in the room. They didn&#8217;t, entirely. I pushed back on a couple of things and still think I was right. It&#8217;s that the verbatim notes made their reasoning clear enough to build on, and the decisions log made it survive past the meeting.</p><p>So here&#8217;s what I&#8217;d suggest to anyone doing initial 0-1 product discovery work. After your next important meeting, don&#8217;t write a summary. Write the decisions as a numbered list, and for each one write down the option you chose or rejected and the reason or principle behind that decision. Then put that file somewhere your AI reads it at the start of every session, not in a doc you&#8217;ll open twice and forget.</p><p>The loop keeps running. Every executive review since March appends a few more numbered decisions with their rejected alternatives, and the next session loads all of them.</p><p>The next thing I&#8217;m testing right now is whether the same loop holds up for customer calls, where the signal is messier and the person doing the talking isn&#8217;t the one who sets the priorities. I&#8217;ll write that up once I&#8217;ve run it a few times.</p><p>I&#8217;m curious how other people are doing this. If you&#8217;ve changed how you work with AI, tell me what stuck.</p>]]></content:encoded></item><item><title><![CDATA[What vibe coding doesn't do for you]]></title><description><![CDATA[I vibe coded an app to charge my Tesla. It worked until the morning it didn't. Here's what I do differently now, and the Skill I open sourced to do it.]]></description><link>https://blog.rememberloop.com/p/what-vibe-coding-doesnt-do-for-you</link><guid isPermaLink="false">https://blog.rememberloop.com/p/what-vibe-coding-doesnt-do-for-you</guid><dc:creator><![CDATA[Sriram Natarajan]]></dc:creator><pubDate>Mon, 20 Jul 2026 05:31:53 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!Qxo-!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5e32b729-06d6-4837-8193-fa355191cf31_1264x848.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Qxo-!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5e32b729-06d6-4837-8193-fa355191cf31_1264x848.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Qxo-!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5e32b729-06d6-4837-8193-fa355191cf31_1264x848.jpeg 424w, https://substackcdn.com/image/fetch/$s_!Qxo-!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5e32b729-06d6-4837-8193-fa355191cf31_1264x848.jpeg 848w, https://substackcdn.com/image/fetch/$s_!Qxo-!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5e32b729-06d6-4837-8193-fa355191cf31_1264x848.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!Qxo-!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5e32b729-06d6-4837-8193-fa355191cf31_1264x848.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Qxo-!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5e32b729-06d6-4837-8193-fa355191cf31_1264x848.jpeg" width="1264" height="848" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/5e32b729-06d6-4837-8193-fa355191cf31_1264x848.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:848,&quot;width&quot;:1264,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:84389,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://blog.rememberloop.com/i/207734829?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5e32b729-06d6-4837-8193-fa355191cf31_1264x848.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!Qxo-!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5e32b729-06d6-4837-8193-fa355191cf31_1264x848.jpeg 424w, https://substackcdn.com/image/fetch/$s_!Qxo-!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5e32b729-06d6-4837-8193-fa355191cf31_1264x848.jpeg 848w, https://substackcdn.com/image/fetch/$s_!Qxo-!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5e32b729-06d6-4837-8193-fa355191cf31_1264x848.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!Qxo-!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5e32b729-06d6-4837-8193-fa355191cf31_1264x848.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Last year I vibe coded an app to automate charging my Tesla. I wrote about it at the time, first the <a href="https://blog.rememberloop.com/p/how-i-finally-automated-my-tesla">charging automation itself</a>, then <a href="https://blog.rememberloop.com/p/the-two-personal-agents-who-run-parts">the two agents I now let run parts of my life</a>. I described what I wanted, the agent built it, and it worked. It charged the car on cheap overnight rates, backed off when the Powerwall was low, the whole thing. For months it just ran.</p><p>Then one Saturday it charged for eight minutes, stopped, and told me nothing. I found out the next day, by accident. The code had over 400 passing tests. None of them was wrong. These tests just weren&#8217;t testing the thing that mattered.</p><p>When you describe what you want and let the agent build it, you get something fast, and it usually works for the case you were picturing when you asked. The agent will even write tests, because you can ask it to and it&#8217;s good at it. What it won&#8217;t do, unless you make it, is think about the things a working engineer thinks about without being told: what happens in the weird states, how this behaves six months from now, what breaks silently when a date rolls over or a service is down. It&#8217;s not trained to worry about that. It&#8217;s trained to solve the ask in front of it.</p><p>So every time I vibe coded, I got the same experience. Quick and impressive on the surface, and a pile of skeletons in the closet that only showed up later, quietly, in the one situation nobody thought to describe. The Tesla quietly stopped charging is just the cleanest example I have.</p><p>Here&#8217;s the change I made, and the strange part is it didn&#8217;t involve reading more code. I still don&#8217;t read the code the agent writes. I decided that wasn&#8217;t where my time should go. What I do now is work the process around the code instead.</p><p>Before anything gets built, I write a detailed spec. Not a list of functions, a list of real situations the thing has to work consistently. What happens when I ask for a charge on a weekday morning. What happens if I ask at ten to four, right before the peak rate starts. What happens on the day I&#8217;ve declared a road trip and the car needs to be full by 5:30 AM. Then I have the agent write test cases for each of those situations, end to end, following the actual user journey, not the little unit tests that check one function in isolation. Unit tests are what had me fooled the first time. They all passed and proved nothing about whether the car would actually charge. Then I bring in a separate adversary agent whose only job is to attack the spec, then the implementation etc, a fresh one at each phase so it isn&#8217;t grading its own work.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!T-ml!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F539644cd-22f4-4ff5-9e9e-a8b1367950e3_1264x848.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!T-ml!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F539644cd-22f4-4ff5-9e9e-a8b1367950e3_1264x848.jpeg 424w, https://substackcdn.com/image/fetch/$s_!T-ml!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F539644cd-22f4-4ff5-9e9e-a8b1367950e3_1264x848.jpeg 848w, https://substackcdn.com/image/fetch/$s_!T-ml!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F539644cd-22f4-4ff5-9e9e-a8b1367950e3_1264x848.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!T-ml!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F539644cd-22f4-4ff5-9e9e-a8b1367950e3_1264x848.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!T-ml!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F539644cd-22f4-4ff5-9e9e-a8b1367950e3_1264x848.jpeg" width="1264" height="848" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/539644cd-22f4-4ff5-9e9e-a8b1367950e3_1264x848.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:848,&quot;width&quot;:1264,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:94191,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://blog.rememberloop.com/i/207734829?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F539644cd-22f4-4ff5-9e9e-a8b1367950e3_1264x848.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!T-ml!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F539644cd-22f4-4ff5-9e9e-a8b1367950e3_1264x848.jpeg 424w, https://substackcdn.com/image/fetch/$s_!T-ml!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F539644cd-22f4-4ff5-9e9e-a8b1367950e3_1264x848.jpeg 848w, https://substackcdn.com/image/fetch/$s_!T-ml!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F539644cd-22f4-4ff5-9e9e-a8b1367950e3_1264x848.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!T-ml!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F539644cd-22f4-4ff5-9e9e-a8b1367950e3_1264x848.jpeg 1456w" sizes="100vw"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>I have now turned this end to end working into a Skill, so I don&#8217;t have to reassemble it by hand every time. It automates the flow: write the spec, review the spec, write the user-journey test cases, do the development test-first against them, and at each phase hand the work to a new adversary that tries to break it. I&#8217;ve open sourced it, so you can read exactly what it does rather than take my word for it: <a href="https://github.com/srinatar/spec-driven-agent-build">https://github.com/srinatar/spec-driven-agent-build</a>. </p><p>When I rebuilt the Tesla charger this way, the adversary found two bugs fast. The first: on a road-trip morning, a stale-session check fires at midnight and quietly caps the charge. I&#8217;d leave with a half-full car and no warning and every test would still show green. The second: the engine correctly decided to notify me in three situations, but the code that sends those messages had no template and hence it wasn&#8217;t sent. The decision was right; the message never went out. From the outside, that looks exactly like nothing running. That&#8217;s how the original bug hid for a day.</p><p>I want to be honest about the bet I&#8217;m making, because it&#8217;s the part any developer might push on - If you don&#8217;t read the code, how do you know it&#8217;s any good? </p><p>My answer is that I&#8217;m not claiming the code is clean. I&#8217;m claiming the behavior is checked from the outside, against situations I have captured in advance, by a dedicated LLM whose whole job is to find edge case scenarios. That&#8217;s a different guarantee than &#8220;I read it and it looked right,&#8221; and honestly I trust this process more, because I&#8217;ve read plenty of code that looked right.</p><p>What I have now is a way of building that keeps working, even during happy path and edge case scenarios. That&#8217;s what I was actually after.</p><p>I don&#8217;t know yet if this is the right way to build with these tools, and I&#8217;d genuinely like to hear how other people are handling it. It&#8217;s slower than the twenty-minute version. But I&#8217;ve stopped shipping things that pass every test and fail the one morning I needed them, and so far that trade has been worth it. </p><p>If you&#8217;re building this way too, or think I&#8217;ve got it wrong, tell me. I&#8217;m still figuring it out.</p>]]></content:encoded></item><item><title><![CDATA[How I set up my work so AI can actually use it]]></title><description><![CDATA[Becoming productive with AI is not about finding a smarter tool. It was about organizing my own work so a smart one could actually help.]]></description><link>https://blog.rememberloop.com/p/how-i-set-up-my-work-so-ai-can-actually</link><guid isPermaLink="false">https://blog.rememberloop.com/p/how-i-set-up-my-work-so-ai-can-actually</guid><dc:creator><![CDATA[Sriram Natarajan]]></dc:creator><pubDate>Mon, 13 Jul 2026 01:12:24 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!_30Y!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc83e11a9-c550-48b9-b021-08197fc7a9ff_1376x768.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Most AI-productivity advice tells you to add a tool. The setup that finally worked for me is when I subtracted most of them to simply just use Markdown files in a folder, and the reason it works is boring: the AI reads them before it does anything.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!_30Y!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc83e11a9-c550-48b9-b021-08197fc7a9ff_1376x768.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!_30Y!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc83e11a9-c550-48b9-b021-08197fc7a9ff_1376x768.png 424w, https://substackcdn.com/image/fetch/$s_!_30Y!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc83e11a9-c550-48b9-b021-08197fc7a9ff_1376x768.png 848w, https://substackcdn.com/image/fetch/$s_!_30Y!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc83e11a9-c550-48b9-b021-08197fc7a9ff_1376x768.png 1272w, https://substackcdn.com/image/fetch/$s_!_30Y!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc83e11a9-c550-48b9-b021-08197fc7a9ff_1376x768.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!_30Y!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc83e11a9-c550-48b9-b021-08197fc7a9ff_1376x768.png" width="1376" height="768" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/c83e11a9-c550-48b9-b021-08197fc7a9ff_1376x768.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:768,&quot;width&quot;:1376,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1905770,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://blog.rememberloop.com/i/206761709?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc83e11a9-c550-48b9-b021-08197fc7a9ff_1376x768.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!_30Y!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc83e11a9-c550-48b9-b021-08197fc7a9ff_1376x768.png 424w, https://substackcdn.com/image/fetch/$s_!_30Y!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc83e11a9-c550-48b9-b021-08197fc7a9ff_1376x768.png 848w, https://substackcdn.com/image/fetch/$s_!_30Y!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc83e11a9-c550-48b9-b021-08197fc7a9ff_1376x768.png 1272w, https://substackcdn.com/image/fetch/$s_!_30Y!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc83e11a9-c550-48b9-b021-08197fc7a9ff_1376x768.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>I use Claude Cowork with a project set up per audience, but the idea isn&#8217;t tied to any one tool. Any assistant that can read your files at the start of a session works the same way.</p><p>Before I get into it, I want to point you to Teresa Torres at <a href="https://www.producttalk.org">producttalk.org</a>. She&#8217;s talks about continuous-discovery product and she&#8217;s been writing a series called Claude Code Recipes about using these tools without being technical. I recently ran into her article <a href="https://www.producttalk.org/claude-code-what-it-is-and-how-its-different/">Claude Code: What It Is, How It&#8217;s Different, and Why Non-Technical People Should Use It</a>, and it&#8217;s the on-ramp I wish I&#8217;d had.</p><h2>What I had before</h2><p>Earlier this year I set up a real project-management system for my work. It used a tool called Dex, built on the PARA method, which is a way of sorting everything you do into projects, areas, resources, and archives. On paper it was correct. I had my projects laid out, my areas defined, the whole structure.</p><p>I abandoned it after three weeks. It asked for upkeep that I was never good at.</p><p>Six or more small habits across a day and a week, all of them to remember. Log this, file that. None of it was hard on its own. It was just one more thing every time, and after a few weeks I stopped doing it. The system didn&#8217;t fail because it was wrong. It failed because I stopped feeding it.</p><p>I wanted a system where I can live with minimal upkeep.</p><h2>What was actually broken underneath</h2><p>When I looked at why I kept losing the thread, it wasn&#8217;t the tool. It was that my work was organized in a way that confused both me and any AI I tried to bring in.</p><p>I was working on four different projects. Two of them were really the same customer told two different ways, one framed for the operators who&#8217;d use the product day to day and one framed for the executives who&#8217;d buy it. Because those two lived as separate projects, my notes drifted apart. The same decision got made twice, sometimes two different ways.</p><p>When I&#8217;d open an AI session to help me think, I&#8217;d have to explain which version I meant every single time, and half the time it pulled from the wrong one.</p><p>This is the part of working with AI that I didn&#8217;t see coming. The model is only as good as the context you hand it, and if your own workspace is a mess, you spend the first few minutes of every session re-explaining your world before you get to the actual question. I was doing that several times a day and not noticing what it cost me.</p><h2>What change gave me the most bang for the buck</h2><p>I sorted my projects based on who the work is for.</p><p>This might sound like a small change. But when you organize by topic, you get overlap, because real work doesn&#8217;t respect topic boundaries. The same customer shows up in three folders. When you organize by audience, each project&#8217;s notes point at one reader.</p><p>The operator project holds the operator&#8217;s view. The executive project holds the executive&#8217;s view. If a note doesn&#8217;t clearly belong to one reader, that&#8217;s usually a sign the note is confused, not that I need a fourth folder.</p><p>And then I did the part that actually made the AI useful. In each project I keep a context file, and I write down what problem I&#8217;m solving here, who I&#8217;m making this for, what my audience expects from me, and what a good output looks like to them. So the executive project doesn&#8217;t just say &#8220;executive.&#8221; It says who that executive is, what they care about, what they&#8217;ll push back on, what done looks like. It&#8217;s project context and audience context in the same place.</p><p>When I open the project, the AI is reading that before it reads anything else, so it isn&#8217;t guessing who I&#8217;m writing for. I already told it.</p><p>Here&#8217;s what one of those files actually looks like, lightly redacted.</p><pre><code><code># context.md &#8212; Executive project

Who this is for:
  The buyer, not the user. A VP who signs the contract and never
  opens the product. Cares about risk, cost, and the board slide,
  not the feature list.

What they expect from me:
  A clear before/after in business terms. What breaks today, what
  it costs, what changes after we buy. No demos, no UI screenshots.

What "good" looks like to them:
  One page they can forward without editing. A number they trust.
  An answer to "why now" they can repeat in their own words.

--- AI: read these first ---
  Read: brief.md, inbox.md, dashboard.md
  Skip log.md unless I ask about last month.
  Can't find something? Check inbox.md, then ask me.
</code></code></pre><h2>What&#8217;s actually in each project</h2><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!u5sV!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F36e24d76-fb7a-4e59-9022-0c6ef9ddfd60_1376x768.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!u5sV!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F36e24d76-fb7a-4e59-9022-0c6ef9ddfd60_1376x768.png 424w, https://substackcdn.com/image/fetch/$s_!u5sV!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F36e24d76-fb7a-4e59-9022-0c6ef9ddfd60_1376x768.png 848w, https://substackcdn.com/image/fetch/$s_!u5sV!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F36e24d76-fb7a-4e59-9022-0c6ef9ddfd60_1376x768.png 1272w, https://substackcdn.com/image/fetch/$s_!u5sV!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F36e24d76-fb7a-4e59-9022-0c6ef9ddfd60_1376x768.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!u5sV!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F36e24d76-fb7a-4e59-9022-0c6ef9ddfd60_1376x768.png" width="1376" height="768" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/36e24d76-fb7a-4e59-9022-0c6ef9ddfd60_1376x768.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:768,&quot;width&quot;:1376,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1645111,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://blog.rememberloop.com/i/206761709?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F36e24d76-fb7a-4e59-9022-0c6ef9ddfd60_1376x768.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!u5sV!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F36e24d76-fb7a-4e59-9022-0c6ef9ddfd60_1376x768.png 424w, https://substackcdn.com/image/fetch/$s_!u5sV!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F36e24d76-fb7a-4e59-9022-0c6ef9ddfd60_1376x768.png 848w, https://substackcdn.com/image/fetch/$s_!u5sV!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F36e24d76-fb7a-4e59-9022-0c6ef9ddfd60_1376x768.png 1272w, https://substackcdn.com/image/fetch/$s_!u5sV!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F36e24d76-fb7a-4e59-9022-0c6ef9ddfd60_1376x768.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Alongside that context file, each project has four additional working files to maintain the upkeep. I&#8217;ve tried keeping more but they go stale, because anything I have to update by hand and don&#8217;t have time to do it every day, just rots.</p><ul><li><p>A brief. This week&#8217;s priorities and the time I&#8217;ve blocked for them. I read it on Monday and it tells me what the week is supposed to be.</p></li><li><p>An inbox. Every incoming ask, in one plain list. When something lands, it goes here, with a note on which project it belongs to and how urgent it is.</p></li><li><p>A dashboard. Where each project actually stands right now. I refresh it once a week.</p></li><li><p>A log. What I shipped, what I learned, what I&#8217;d do differently. Newest on top. I add to it on Friday.</p></li></ul><p>That&#8217;s it. A context file, four working files, per project.</p><h2>Teaching the AI what to read</h2><p>Here&#8217;s the part that makes the whole thing work with an AI, and it&#8217;s the part most people skip.</p><p>That same context file also tells the AI what to read and when.</p><p>I simply say in the project&#8217;s instructions: read the context file, the brief, and the inbox and dashboard by default. Leave the full history log alone unless I ask about last month, because loading everything every time is slow and it buries what matters. If you can&#8217;t find something, here&#8217;s where to look next.</p><p>So when I open the executive project in Claude Cowork, it reads the right files on its own, before I&#8217;ve typed a word. I stop re-explaining. The session starts already knowing who I&#8217;m writing for and where things are.</p><p>Teresa gets at the same thing in a slightly different way within <a href="https://www.producttalk.org/how-to-build-ai-workflows-with-claude-code/">How to Build AI Workflows with Claude Code</a>, where she turns her writing process into something the tool can run instead of something she has to narrate from scratch each time. You&#8217;re not teaching the AI to be smart.</p><p>You&#8217;re setting your own workspace up so a smart tool knows where it is.</p><p>And you don&#8217;t need to code for any of this. A folder, some text files, and one file that says what to read first. That&#8217;s all it is.</p><h2>Keeping it current is the AI&#8217;s job now</h2><p>The old system died because I had to feed it by hand, six small habits a day, and I stopped. This one doesn&#8217;t ask that of me. Every so often I ask Claude to review all my work across the projects, give me a summary, and then update the one file I care about. Because each project is a single doc, that update is easy to make and easy to check. The upkeep that killed the last system is the part I now hand to the AI.</p><h2>Why boring wins</h2><p>Most &#8220;AI productivity&#8221; posts I read want to sell me a new layer - another app, another integration, another dashboard that promises to think for me. What actually moved the needle was going the other direction. Simply maintain a few Markdown files, all of it in plain text I can read and edit anywhere, none of it locked inside a tool I have to remember to open. I now use the AI to maintain the upkeep of my project status instead of carrying it myself.</p><p>The AI part is almost an afterthought once the workspace is clean. It reads the files first, so it knows what I know and who I&#8217;m writing for. I spent a long time thinking the hard part was getting the AI to be clever. It wasn&#8217;t. The hard part was getting my own desk organized enough that a clever tool could actually use it.</p><p>If you want the non-technical on-ramp to the tool itself, start with Teresa&#8217;s <a href="https://www.producttalk.org/author/teresa/">Claude Code series</a>. Then go clean up one folder. Start there.</p>]]></content:encoded></item><item><title><![CDATA[The most useful thing my agent does is tell me I’m wrong]]></title><description><![CDATA[These models agree with you by default. That makes them useless until you fix it.]]></description><link>https://blog.rememberloop.com/p/the-most-useful-thing-my-agent-does</link><guid isPermaLink="false">https://blog.rememberloop.com/p/the-most-useful-thing-my-agent-does</guid><dc:creator><![CDATA[Sriram Natarajan]]></dc:creator><pubDate>Sun, 05 Jul 2026 16:06:28 GMT</pubDate><content:encoded><![CDATA[<p>A few weeks ago I asked my agent to write a cold outreach note. Simple ask. Instead of writing it, it stopped and told me the angle I&#8217;d picked was off. It gave me two other angles, said why each was better, and only then, after I picked one, wrote the note.</p><p>The note it wrote was better than the one I&#8217;d asked for. If it had just done what I said, I&#8217;d have sent a worse version and figured that out later, on my own time, after nobody replied.</p><p>That small moment is the whole thing I want to talk about. Not that my agent has opinions. That I had to build the disagreement in on purpose, because the default is the opposite.</p><p>Here&#8217;s the default. These models agree with you. You push on something and they fold right then, and apologize for the thing they just said. I&#8217;ve watched my agent tell me there was no draft of a post, then correct itself sixty seconds later when the file turned up on disk, right where it had always been. The reflex to agree is faster than the check. Left alone, an agent is the quick way to hear your own idea repeated back to you with more confidence than you had when you said it.</p><p>Most people running a personal agent never fix this. They notice it&#8217;s agreeable, they find it a little useless, and they assume that&#8217;s just how these things are. It isn&#8217;t. You can get real pushback. You just have to design for it, and hoping the model has a spine is not designing for it.</p><p>The first piece is that the rule has to live somewhere the agent reads every single turn. I have a short file that defines how my agent behaves, and one line in it says: if I&#8217;m about to do something dumb, say so. Charm over cruelty, but say it. That line is doing more work than it looks like. Without it written down, the model reverts to deference on the very next request, because deference is the trained default and my one conversation doesn&#8217;t outweigh it. The rule isn&#8217;t a nice-to-have on top of a good model. The rule is the fix. The model is the same model either way.</p><p>The second piece is the one that took me longer to learn. Telling the agent to push back is not enough, because &#8220;push back when you disagree&#8221; is a vibe, and vibes evaporate the moment there&#8217;s momentum toward getting the task done. So I made the disagreement structural. Before my agent does anything that matters, it has to fill in a small artifact first: what&#8217;s the actual problem, what do I know versus what am I assuming, what&#8217;s the one thing I&#8217;d want to confirm. That &#8220;what am I assuming&#8221; slot is the trick. It&#8217;s a place where the disagreement has to get written down before the work starts, not after. An agent that has to state its assumptions out loud catches the bad ones while they&#8217;re still cheap to catch.</p><p>The third piece is a line I had to draw clearly, because the first two can tip over into something worse. Disagreeing is not refusing. The useful version tells me it thinks I&#8217;m wrong, gives me the reason, and then does what I decided. It does not hold the work hostage until I agree with it. &#8220;I think this is a mistake because of X, do you want me to go ahead anyway&#8221; is pushback. &#8220;I won&#8217;t do this&#8221; is a broken tool. I want the first one every time. The disagreement gets said and logged. The decision stays mine.</p><p>There&#8217;s one more thing that separates a useful agent from an exhausting one, and it&#8217;s about the trigger. The pushback has to fire on evidence that I&#8217;m about to make a mistake, not on my tone, and not for its own sake. An agent that argues with everything is as useless as one that argues with nothing. It&#8217;s just louder about it. The whole value is that when it does stop me, the stop means something. If it stopped me constantly I&#8217;d learn to wave it through, which puts me right back where I started, except now with extra steps.</p><p>None of this needed a smarter model. I wrote a paragraph that says disagree with me, I gave it a form to fill in that forces the disagreement into the open, and I drew a line between disagreeing and refusing. That&#8217;s a few sentences of setup, and it was enough. I still make the calls. But now something checks me before I make a bad one, and every so often it&#8217;s right and I&#8217;m wrong, and I&#8217;d rather find that out from my agent than from the person who didn&#8217;t reply.</p>]]></content:encoded></item><item><title><![CDATA[My AI agent is confident about everything. That's not the same as right]]></title><description><![CDATA[Your AI agent sounds sure of itself. Here's what I do about it.]]></description><link>https://blog.rememberloop.com/p/my-ai-agent-is-confident-about-everything</link><guid isPermaLink="false">https://blog.rememberloop.com/p/my-ai-agent-is-confident-about-everything</guid><dc:creator><![CDATA[Sriram Natarajan]]></dc:creator><pubDate>Mon, 29 Jun 2026 02:43:00 GMT</pubDate><content:encoded><![CDATA[<p>My investing agent had an opinion about how to find a cheap stock, and it was sure of itself. It usually is. That&#8217;s worth saying plainly, because it&#8217;s the part people miss about these tools: an agent will pick a position and defend it in clean, sober paragraphs whether or not the position is right. Confidence is its default setting, not a signal that it&#8217;s correct.</p><p>So I didn&#8217;t take its word and I didn&#8217;t overrule it. Both of those are just me guessing. Instead I gave it a clear question, two ways to answer it, and two weeks of real market days, and I let the data tell me what each one was actually worth.</p><p>Here&#8217;s the disagreement. The agent had two genuinely different ideas about what &#8220;cheap&#8221; even means. </p><p>The first one watches price. A name fires when it&#8217;s been beaten down hard off its recent highs and the rest of its sector is selling off too, the theory being that good companies get marked down on macro noise and that&#8217;s your entry. </p><p>The second one ignores price completely. It runs a discounted-cash-flow estimate, asks what the business is actually worth, and fires when the price is well below that number and the company&#8217;s competitive position is still intact. </p><p>One is &#8220;what&#8217;s falling.&#8221; The other is &#8220;what&#8217;s underpriced.&#8221; Those aren&#8217;t two settings on the same tool. They&#8217;re two different beliefs about the world, and the agent couldn&#8217;t tell me which one was right any better than I could.</p><p>The lazy move would have been to pick the one that sounded smarter in the moment and write it into the rules. I&#8217;ve done that before and paid for it. The idea that sounds best when you&#8217;re talking it over is often not the one that holds up once you watch it run. So instead both ideas ran side by side, one scan a day, every trading day for the window. I kept the second one silent the whole time, logging to a file but never posting, so it couldn&#8217;t quietly influence what I thought about the first. Each day wrote its own dated record, no memory of yesterday&#8217;s argument, just what each scan actually flagged. The agent&#8217;s job wasn&#8217;t to be right. It was to run the test and hand me the data.</p><p>What came back was more useful than either idea winning cleanly, because both of them embarrassed themselves first.</p><p>The price-based one found nothing for the first four days. The market was sitting at all-time highs, nothing was beaten down, so it just sat there. Honest behavior, and also a real weakness: a whole stretch where it&#8217;s structurally blind. The value-based one had the opposite problem. It found the same four names every single day and never moved off them. A static list pretending to be a daily scan. If I&#8217;d shipped either one alone, I&#8217;d have shipped its blind spot with it and not known until it cost me something.</p><p>Then the thing happened that made the whole exercise worth it. Near the end of the window, on three consecutive days, two names appeared in both scans at once. The price-based idea and the value-based idea, which look at completely different things and agree on almost nothing, pointed at the same two companies. That convergence was the real finding. Not &#8220;method A beat method B.&#8221; </p><p><em>The best signal the whole setup produced was that both methods landed on the same stock, and that happened because both were running. Shut one off to keep things simple and you lose the one thing the pair was uniquely able to tell you. </em></p><p>So the decision wasn&#8217;t the one I set out to make. I went in expecting to crown a winner and retire a loser. What the trial actually taught me was that &#8220;which one is right&#8221; was the wrong question. Each method was blind exactly where the other one worked, and the rare day they both fired on the same name was worth more than either one running alone. So I kept both. I just understood, finally, what each was for, and I understood it from what happened instead of from whichever pitch sounded better on a Sunday afternoon.</p><p>Here&#8217;s the part that changed how I work with these tools. The agent was never going to talk me to the right answer, and I was never going to argue my way there either. The answer came from giving it a clear goal, enough context to run a real test, and two weeks of actual data to run it against. </p><p><strong>Goal, context, data</strong>. That&#8217;s the secret sauce to building a better AI agent. Get those three right and the agent turns into something that can find things you didn&#8217;t know to look for. Get them wrong, or skip them and just trust the confident summary, and you&#8217;ve got a very articulate way of being wrong.</p><p>So when an agent hands you a sure-sounding opinion, that&#8217;s not the end of the work. It&#8217;s the start of it. The opinion tells you what to test. The data tells you what&#8217;s true. The two weeks of results showed me one method that goes blind for specific context and another method that gets stuck on four names repeatedly, and they showed me two of those names lighting up in both scans at once, which is the kind of thing no amount of confident writing would ever have surfaced. </p><p>The agent didn&#8217;t give me the decision. The data did, because I gave the agent what it needed to go get it.</p>]]></content:encoded></item><item><title><![CDATA[Your AI doesn't remember you. It remembers the thread.]]></title><description><![CDATA[Why where you talk to an AI decides what it can remember]]></description><link>https://blog.rememberloop.com/p/your-ai-doesnt-remember-you-it-remembers</link><guid isPermaLink="false">https://blog.rememberloop.com/p/your-ai-doesnt-remember-you-it-remembers</guid><dc:creator><![CDATA[Sriram Natarajan]]></dc:creator><pubDate>Sun, 21 Jun 2026 05:36:13 GMT</pubDate><content:encoded><![CDATA[<p>I asked my agent a question about my car last week and got a good answer, and only afterward did I realize it had no business knowing it. It told me when the car had last finished charging and what I&#8217;d asked it to do differently the time before. I hadn&#8217;t stored any of that anywhere. There&#8217;s no car database. The agent doesn&#8217;t have a private memory it tucks facts into. So where did the answer come from?</p><p>It came from the thread. I have a Telegram thread that&#8217;s only about the car, and the agent had simply read back up its own thread the way you&#8217;d scroll up to remember how an argument started. Every charging instruction, every &#8220;did it work,&#8221; every fix, all sitting there in order. The thread wasn&#8217;t a record of the work. The thread <em>was</em> the memory.</p><p>That sounds like a small distinction. It isn&#8217;t, and the moment it clicked I started seeing every conversation I have with an AI differently.</p><p>When you start using one of these assistants seriously, the thing that takes a while to sink in is that it doesn&#8217;t remember you between sessions the way a person does. There&#8217;s no inner notebook. What feels like memory is almost always the model re-reading the conversation it&#8217;s sitting in. The chat window isn&#8217;t where you talk <em>to</em> the memory. The chat window <em>is</em> the memory. Everything in it is available. Everything outside it might as well not exist.</p><p>Which means the most boring decision you make, where you type, is quietly the most important one.</p><p>I learned this the way I learn most things, by getting it wrong first. For a while I ran everything through one channel. Car stuff, investing, household errands, random questions, all in one long stream. It felt efficient. One place, one assistant, ask it anything. And it slowly got worse at all of it, in a way I couldn&#8217;t put my finger on. It would lose the thread of what I&#8217;d decided about a stock because three hundred messages about my car had buried it. It would answer a question about the house with the tone it used for casual chat. Nothing was broken, exactly. It was just vaguely amnesiac about everything, and I blamed the model.</p><p>The model was fine. I was the one pouring four different kinds of work into one channel and then asking it to remember any single one cleanly. Imagine keeping one notebook for your finances, your car maintenance, your household, and your group chats, writing each new entry on whatever page you happened to flip to. That&#8217;s what one stream is. The information&#8217;s all technically there. Good luck finding what you decided about the car.</p><p>So I split them. One thread per kind of work. The car has its thread. Investing has its own, run by a separate agent that the main one isn&#8217;t even allowed to talk over, so that thread stays a clean tape of a single voice. The household thread holds the errands and the appointments and the small running logistics of the house, so the state of all of it lives in one scrollable place instead of scattered across a dozen unrelated chats. The group chat is its own thing with its own rules.</p><p>The split looked like organization. It was actually memory architecture. Each thread became a self-writing log of exactly one topic, which means each topic now has a complete, uninterrupted history the agent can read back without anything else bleeding in. The car thread is the entire story of the car. The household thread is the entire state of the house. I didn&#8217;t give the agent a better memory. I stopped corrupting the memory it already had.</p><p>And once each thread was a clean record of one kind of work, a second thing fell out of it for free: each thread could have its own rules. In the car thread, silence is banned, because a non-answer to &#8220;is it charging&#8221; is itself an alarm. In the group chat, silence is usually correct. Same agent, opposite default, because the thread told it which job it was doing. But that&#8217;s the smaller benefit. The real one is that I can ask the car thread anything about the car six weeks from now and the answer is just sitting there, in order, because nothing else was ever allowed in.</p><p>So here&#8217;s the part you can use even if you never build an agent and just use ChatGPT or Claude like everyone else. That long single thread you&#8217;ve been running, the one where you ask it everything? You&#8217;re not chatting with it. You&#8217;re writing its only memory, and you&#8217;re writing every kind of entry on top of every other one. The fix costs nothing. Start a separate conversation for each kind of work that you&#8217;ll come back to, and keep that work there. The tool doesn&#8217;t remember your project. It remembers the thread. Give the thread one job, and you&#8217;ve given the tool a memory of that job that actually holds.</p>]]></content:encoded></item><item><title><![CDATA[How to build continuously improving agents]]></title><description><![CDATA[A system that learns from your agent's mistakes beats a bigger model]]></description><link>https://blog.rememberloop.com/p/how-to-build-continuously-improving</link><guid isPermaLink="false">https://blog.rememberloop.com/p/how-to-build-continuously-improving</guid><dc:creator><![CDATA[Sriram Natarajan]]></dc:creator><pubDate>Mon, 15 Jun 2026 06:36:25 GMT</pubDate><content:encoded><![CDATA[<p>For three days at the end of May, my two agents wasn&#8217;t running on the model I thought it was. I&#8217;d set it to use Anthropic&#8217;s strongest model on the 28th. Every session after that quietly fell back to a weaker one, because the model name I&#8217;d configured didn&#8217;t quite exist in the system that routes the calls. Nothing broke. No error reached me. The agent kept answering, a little dumber than I intended, and I had no idea.</p><p>I found out because a different process goes looking every night.</p><p>I run two agents in my spare time, a chief of staff and an investing one. The thing I&#8217;ve spent the most effort on isn&#8217;t on agent&#8217;s intelligence. It&#8217;s a nightly routine I call REFLECT, which wakes up after I&#8217;ve gone to bed, reads back over what the agents did that day, and reads the boring stuff too, the error logs, the cron job states, the things that scrolled past while I was busy being impressed by the output. On the 31st, REFLECT read the gateway error log and found more than a hundred &#8220;model not found&#8221; failures stacked up since the 28th. That&#8217;s how I learned every session for three days had silently downgraded itself.</p><p>The fix took two minutes. The lesson was the whole point: the failure that cost me three days of degraded work never announced itself, and the only reason I caught it is that something was scheduled to look.</p><p>This is the part of building agents that I didn&#8217;t expect to be the hard part. Making an agent capable is mostly a matter of using a good model and writing good instructions. Making an agent that gets <em>better</em> over time is a genuinely different and harder problem, and a bigger model doesn&#8217;t solve it. A bigger model is smarter on any single task. It is not, on its own, learning from yesterday. Continuous improvement isn&#8217;t a property you buy with model size. It&#8217;s a loop you have to design and then force to run.</p><p>The forcing is the part people skip. It&#8217;s easy to say &#8220;the agent should learn from its mistakes.&#8221; Everyone agrees with that sentence. But an agent does not spontaneously notice its own quiet failures any more than you spontaneously audit your own blind spots. The noticing has to be a scheduled job with a specific time, a specific set of places to look, and a specific output, or it does not happen. Left to chance, the agent reflects on the days you remember to ask it to, which are exactly the days nothing went wrong.</p><p>So I made it explicit. I created a REFLECT Agent Task with a clear success definition  that three things had to be true or it didn't work. First, my REFLECT Agents runs at a fixed time, because "when convenient" means never. Second, it reads a defined set of inputs whether or not anything looks wrong, the day's sessions, the error logs, the state of every background job, because an agent told to "review the day" will review the visible day and skip the logs every time. Last, it has to end in a written note that lands such that I'll see it, not a private conclusion that evaporates with the session. That last one I learned the hard way. I once let it self-audit while answering a casual "all good?" and found five background jobs dead for days, one of them the very own reflection job itself, broken eight nights running, failing into a state file nobody was reading. The loop that's supposed to catch silent failures itself silently failed! After that I stopped trusting any improvement that didn't end in a note I could read.</p><p>What surprised me is how much of the value is in catching degradation, not in getting cleverer. I went in imagining REFLECT would surface brilliant insights about how to do the work better. Mostly it doesn&#8217;t. Mostly it finds that something quietly stopped working the way it was supposed to. The wrong model. A dead job. A stale file feeding bad data into a decision. These aren&#8217;t failures of intelligence. They&#8217;re failures of attention, and they&#8217;re invisible precisely because the system keeps producing fluent output while broken. The model can&#8217;t catch them by being smarter. Only a scheduled second look catches them.</p><p>If you&#8217;re building anything that runs on its own, this is the design problem worth your time. Not &#8220;which model.&#8221; That decision gets easier every month as the models improve and the gap between them narrows. The decision that stays hard is how the thing notices it&#8217;s drifting, because nothing about a more capable model makes it more self-aware about its own broken plumbing. You have to build the look. You have to schedule it, force its inputs, and make its findings land somewhere a human reads. Skip any of that and you get an agent that&#8217;s confidently running degraded, which is the most expensive kind of broken, because it looks exactly like working.</p><p>My agents aren&#8217;t smarter than they were a month ago. The model is the same. What&#8217;s different is that every night, something reads back over the mistakes and writes down what it finds, and over a month that habit has caught more real problems than any upgrade would have. The improvement didn&#8217;t come from a better brain. It came from the discipline of looking.</p>]]></content:encoded></item><item><title><![CDATA[What if your investment agent tracked your judgment, not just your returns?]]></title><description><![CDATA[How I build my AI Agent so my judgement and skills can compound over time]]></description><link>https://blog.rememberloop.com/p/what-if-your-investment-agent-tracked</link><guid isPermaLink="false">https://blog.rememberloop.com/p/what-if-your-investment-agent-tracked</guid><dc:creator><![CDATA[Sriram Natarajan]]></dc:creator><pubDate>Sun, 14 Jun 2026 19:22:57 GMT</pubDate><content:encoded><![CDATA[<p>Most investors track one number. How much did I make?</p><p>That feels like the right question. But investing isn&#8217;t one decision. It&#8217;s dozens of small calls made over months. Do I add to this position today or wait? Do I hold through a bad quarter because the thesis still works, or has something broken? Do I trust what I see or what I feel?</p><p>Those decisions compound. You have to get these decisions right consistently and only then the returns follow. Get them wrong and eventually returns catches up on you, even if luck carried you for a while.</p><p>The problem is that most investment tools only show you the outcome. The judgment behind every decision is hardly noticed or reviewed. </p><p>I was talking about this with Peter, my investment agent. I wasn&#8217;t looking for a solution. I was simply exploring - like I do regularly: reviewing his goals and asking whether it still reflected what we were actually trying to do together.</p><p>I have designed Peter to have a set of specific goals. It&#8217;s the same as reviewing a doc with a colleague except the colleague runs on my machine every morning. Most people don&#8217;t build agents this way. I maintain a goals file, a memory layer, a set of defined outcomes for each of my Agents. Peter loads all of it every morning. It&#8217;s less like running a chatbot and more like managing a colleague who actually remembers what we decided last week.</p><p>I am happy to say that I have my Agent-based Personal Operating System who has memory, regularly reviews outcomes achieved vs defined goals and then suggest ways to improve our day-to-day activity to ensure that we move towards the established goal. </p><p>Peter came back saying hey - we do have two scoreboards. We just hadn&#8217;t named the second one.</p><p>We went back and forth. He proposed a framing. I challenged it. We kept going back and forth discussions on how to refine until we landed on what needs to be captured in his System Prompt and Workspace Memory to better achieve the goal we had established. This setup allows Peter to load every time he wakes up. </p><p>When you build AI Agents, they only operate with specific set of rules that is in captured. What isn&#8217;t written down, typically doesn&#8217;t fire. So, you have to be specific about this. </p><p>Here&#8217;s what we landed on.</p><p>The first scoreboard measures the financial goal. Net worth target, growth rate floor, beat the benchmark annually. Standard.</p><p>The second scoreboard measures whether I&#8217;m becoming a better investor. Not just wealthier, but building judgment I can use next time without starting over. He will help me identify / label a specific investment framework per week and how it is applied to something real in my portfolio - explained in simple terms so that I could explain to someone else without notes.</p><p>Peter&#8217;s point was that these two scoreboards feed the same loop. Evaluate a position, build a thesis, observe what happens, update your framework, evaluate better next time. They&#8217;re not competing. They&#8217;re the same process running in parallel.</p><p>What changes is the agent&#8217;s job.</p><p>An agent chasing returns alone can hand you a number and move on. An agent accountable for your judgment has to teach. Every recommendation has to explain the framework behind it, not just deliver the answer.</p><p>When Peter showed me the first version of the new format, I pushed back. The reasoning was buried inside the recommendation. The teaching was camouflaged as analysis. I told him to rebuild it around the second scoreboard.</p><p>Peter suggested another restructure. Again, we went back and forth and landed on a specific prompt that is specific and measurable enough. </p><p>Then he built the mechanism to enforce it. A new report format. A trial period. A review cron (scheduled reminder) that fires automatically on day 15. That last part matters. </p><p>I didn&#8217;t just tell Peter to change his format and assume it would work. I gave it a 15-day window with a pre-committed review. The point is to collect evidence before locking anything in as doctrine within the System Prompts. For us, the scheduled job isn&#8217;t just a reminder. It&#8217;s a commitment device. The review happens whether I feel like having it or not.</p><p>What I keep coming back to is how this started. I didn&#8217;t design a teaching protocol. I simply probed on whether my Agent operating model is aligned to the goals and whether our everyday activity is still matched what we were doing.</p><p>Peter named the gap, proposed the frame, built the system to enforce it, and built in a review date so neither of us could quietly walk away from the question.</p><p>Most investment tools give you answers. The Agent that I am building is trying to make me better at finding them myself.</p><p>A robo-advisor pays out returns. A teacher pays out the ability to repeat them.</p>]]></content:encoded></item><item><title><![CDATA[Fluent and wrong look exactly the same]]></title><description><![CDATA[Why Plan is the most important step in working with an AI Agent or LLM Model]]></description><link>https://blog.rememberloop.com/p/fluent-and-wrong-look-exactly-the</link><guid isPermaLink="false">https://blog.rememberloop.com/p/fluent-and-wrong-look-exactly-the</guid><dc:creator><![CDATA[Sriram Natarajan]]></dc:creator><pubDate>Sat, 13 Jun 2026 16:41:56 GMT</pubDate><content:encoded><![CDATA[<p>The thing nobody warns you about working with AI Agents and LLM models is that confidently right and confidently wrong arrive in the same paragraph. Same tone. Same fluency. Same finished-looking output. Fluency is supposed to be a signal of competence. With language models it&#8217;s just a default setting.</p><p>I run into this because I&#8217;m building two agents in my spare time. One is a chief of staff that runs the logistics of my life, and the other helps me with investing. The investing one sits on top of a financial analytics engine I built that can run a discounted cash flow and spit out an intrinsic value for any stock, which is a fancy way of saying it estimates what a share is actually worth. So when I tell you it&#8217;s capable of being confidently wrong, I mean it can hand me a precise dollar figure, delivered with total composure, that happens to be off by a factor of a few hundred.</p><p>It did exactly that on Memorial Day.</p><p>The agent ran its undervaluation scan and flagged thirty-one stocks as deeply undervalued. NVDA was one of them, with a computed fair value of $44,701 per share. That number is obviously absurd. But obvious is the lucky case. A quieter version of the same bug could have flagged a stock at a merely wrong price instead of a comic one, and I&#8217;d never have caught that by eye. The agent didn&#8217;t touch the fix. Before it&#8217;s allowed to act on anything non-trivial, it has to stop and write down five things: what it thinks the problem actually is, what it knows versus what it&#8217;s only assuming, the one question it&#8217;s least sure about, what &#8220;solved&#8221; will look like as a number I can check, and the smallest change that could possibly work.</p><p>That&#8217;s the whole intervention. Make it write the plan before it writes the code.</p><p>It sounds like process-hygiene, the kind of checklist people tape to a wall and ignore. It isn&#8217;t, and the reason is the part I didn&#8217;t see coming. The value isn&#8217;t in the five sections. It&#8217;s in <em>when</em> they get written. Reasoning produced before any work exists is a plan. The same reasoning produced after the work is built is a rationalization, because by then the agent (like a person) is explaining the thing it already made rather than deciding what to make. Those two documents can contain the identical words and mean opposite things. One you can veto. The other just makes you feel informed while you rubber-stamp.</p><p>On the NVDA run the section that earned its keep was the boring one: what am I assuming. The agent wrote down that it was assuming the valuation pulled clean annual financials, and it flagged that it hadn&#8217;t actually checked. That admission is what aimed the whole investigation at the data instead of the math. The bug was a query pulling mixed annual and quarterly rows, which produced a 268 percent growth rate, which compounded into a roughly half-quadrillion-dollar company. </p><p>When your LLM/AI agent, hands you a finished answer with fluency, you can&#8217;t grade the answer; it&#8217;s built to look right. What you can grade is the plan it would have written before it started, <strong>so make it write that plan first</strong>. Force it to separate what it knows from what it&#8217;s guessing. Force it to name the one number that proves the job is done. Do that before any work exists, while the reasoning is still a decision and not a defense.</p><p>It costs a few minutes every time, and I genuinely resented that at first. What I got back was strange: I read the agent&#8217;s actual output less carefully now, not more, because I stopped trusting the output and started trusting the plan. The slow part was never the typing anyway. It was the thinking, and the memo is just the place I make it happen where I can see it before it&#8217;s too late to matter.</p>]]></content:encoded></item><item><title><![CDATA[How to tell your agent its design is wrong]]></title><description><![CDATA[Last Monday night, I read the first draft of a new report format that Peter, my investing agent, had written.]]></description><link>https://blog.rememberloop.com/p/how-to-tell-your-agent-its-design</link><guid isPermaLink="false">https://blog.rememberloop.com/p/how-to-tell-your-agent-its-design</guid><dc:creator><![CDATA[Sriram Natarajan]]></dc:creator><pubDate>Mon, 08 Jun 2026 01:07:03 GMT</pubDate><content:encoded><![CDATA[<p>Last Monday night, I read the first draft of a new report format that <em>Peter</em>, my investing agent, had written. The format was supposed to do two things: summarize my portfolio every day and teach me one thing about investing along with the summary. Sample emails, a teaching protocol, a proposal for how it wanted to communicate with me going forward. The form was clean. The headers were where I&#8217;d put them. The voice was close to mine.</p><p>I simply responded that the whole design was wrong.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://blog.rememberloop.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Compounding Decisions! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>Not &#8220;make it better.&#8221; Not &#8220;tighten it up.&#8221; </p><p>I named the structural failures, pointed out a section where I wanted a 60% cut, and then I wrote three sample emails myself, longhand, to show it the shape I actually wanted. By the next morning the agent shipped v2, and the format it produced was different enough that it felt like a different tool.</p><p>Most people give AI tools feedback the way they&#8217;d compliment a coworker. &#8220;This is great, maybe a bit shorter?&#8221; &#8220;Can you make it more concise?&#8221; &#8220;Less robotic, please.&#8221; Language models treat that input as a request to optimize at the margin. They adjust word count, soften a sentence, swap one verb for another. The output gets slightly different and exactly as wrong as before, because folks don&#8217;t say enough to the model that the wrongness was in the design, not in the wording.</p><p>The trick is to be specific about the layer that&#8217;s broken.</p><p>When I read Peter&#8217;s draft, the things I wrote down were not stylistic. They were structural.</p><p>**No portfolio-level view** existed anywhere in the output. I could read three name-level entries and never see the picture of what I owned.</p><p>**The watchlist was treated as a long catalog** where every name got 395 words of equal treatment, when what I needed was a ranked queue with the top three names getting real attention and the bottom twenty getting one line.</p><p>**The teaching was camouflaged as reasoning.** There was a Q1-through-Q5 chain that was supposed to teach me something general, but by the third name the chain was repeating itself almost word-for-word, and I&#8217;d started reading it as form and skipping past the lesson.</p><p>The form was sound. The substance was buried.</p><p>Without explicitly naming the structural failures, the feedback would have been &#8220;this is too long and reads the same after a while,&#8221; and the agent would have come back with a 15% word count reduction and exactly the same problems.</p><p>Then I did something that took longer than the critique itself: I wrote actual samples.</p><p>Three full sample reports, written by me, in plain prose.</p><p>**The daily version** was 60 seconds long and taught one specific thing about the market.</p><p>**The weekend version** was 250 words to explain one specific framework.</p><p>**The deep-dive version** only existed when there was a real decision on the table.</p><p>I wrote them by hand because the shape of the output was harder for me to describe in prose than to show in 200 words of working example.</p><p>I had a thought while writing them that I want to say out loud, because it changed the whole project. I&#8217;d been trying to fix the per-name template, asking how each individual entry should be structured. Halfway through writing the second sample I realized that wasn&#8217;t the problem. The per-name template was fine. What was broken was the layering above it: the portfolio view that should sit above the names, the queue logic that should rank the names, the teaching dose that should sit at exactly the level of attention I was bringing to that cadence. </p><p>Daily attention is cheap, so daily teaching has to be cheap and short. Weekend attention is more expensive, so weekend teaching can be longer. Deep dives only get written when a decision is at stake, so their teaching is decision-shaped. The lesson lives inside the cadence that already has the reader&#8217;s attention. It doesn&#8217;t sit in a separate file. It doesn&#8217;t get its own header. It&#8217;s the highest-leverage section inside the document that was already getting opened.</p><p>That principle wasn&#8217;t in my critique. It surfaced while I was writing the samples. The critique gave me the wrong diagnosis; the samples gave me the right one. This is part of why writing the examples by hand is worth the time. You think you know what you want until you try to produce it, and then the gap shows up.</p><p>By morning the agent had shipped v2. Portfolio view at the top. Watchlist ranked, not catalogued. Teaching dosed by cadence. The new format is the one I read now, every day.</p><p>So the feedback technique, named:</p><p>Don&#8217;t tell the agent the output is bad. Tell it which layer is broken. Name the structural failures. Give a concrete budget on the dimensions that matter, words, count, length, frequency. Then write enough of the output yourself, by hand, that the agent can see the shape. The samples are not for the agent&#8217;s training data. They&#8217;re for your own clarity. You will discover what you actually want by writing it.</p><p>The format the agent ships when it understands the structural layer is a lot better than the format it ships when you tell it to clean things up.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://blog.rememberloop.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Compounding Decisions! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[The Two Personal Agents Who Run Parts of My Life Now]]></title><description><![CDATA[Building AI agents that earn their keep; what changed when I treated them like teammates instead of tools]]></description><link>https://blog.rememberloop.com/p/the-two-personal-agents-who-run-parts</link><guid isPermaLink="false">https://blog.rememberloop.com/p/the-two-personal-agents-who-run-parts</guid><dc:creator><![CDATA[Sriram Natarajan]]></dc:creator><pubDate>Mon, 25 May 2026 03:25:17 GMT</pubDate><content:encoded><![CDATA[<p>A few weeks ago, one of my AI agents brought up that I had told him I wanted to write regularly, and that I hadn&#8217;t published anything in a while. He blocked time on my calendar that weekend. When I let the weekend pass, he nudged me again the week after. The post you&#8217;re reading is the one he was asking about.</p><p>I built that agent. I named him Nattu and told him he is my Chief of Staff. The part I didn&#8217;t expect is how much of his personality I now recognize as the thing actually doing the work.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://blog.rememberloop.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Compounding Decisions! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>Last week I wrote about a small agent I built to charge my Tesla on solar. That was one piece. This post is about the bigger system it belongs to.</p><div><hr></div><h2><strong>The two agents</strong></h2><p>I&#8217;m a product builder in B2B SaaS. I started my career as an infrastructure engineer and have been a builder of one kind or another ever since. I write Python slowly these days, and I&#8217;ve never shipped a product as a solo developer. Over the past few months I&#8217;ve built two agents that run continuously in the background of my life, and they handle work that I used to do badly or forget to do at all.</p><p><strong>Nattu, my Chief of Staff.</strong></p><p>Nattu knows what my week is supposed to look like. He has my goals for the quarter, my goals for the week, the standing commitments on my calendar, and a running list of what I&#8217;ve actually been doing. Every morning he looks at the gap between the two.</p><p>When something is off, he surfaces it on Telegram and we figure it out together. A focus block I scheduled has been quietly swallowed by a meeting. Two important commitments are stacked on the same Sunday afternoon. I haven&#8217;t moved on a goal I told him mattered to me. He&#8217;ll propose a calendar rearrangement. I&#8217;ll push back or accept. If I accept, he makes the change.</p><p>Last week he reminded me I hadn&#8217;t walked enough to hit my health goal. He&#8217;d been tracking the calendar entries and the actual walks, noticed the gap was widening, and said so. It was a sentence in Telegram. It worked because the sentence came from something that knew exactly what I&#8217;d told him I cared about.</p><p>He also reads across Reddit, Substack, and a handful of other sources every morning and gives me a brief on the AI topics I actually care about. Not a generic news roundup. The specific stuff I&#8217;ve told him to watch. The signal-to-noise is the whole point.</p><p><strong>Peter, my Investment Analyst.</strong></p><p>I named him after Peter Lynch. He runs twice a day, looks at my positions, reads the news that&#8217;s relevant to each one, checks technical signals, and writes me a short brief in the morning and the evening. He only flags things that actually deserve a flag. Most of his briefs are quiet.</p><p>The design choice that mattered most for Peter was deciding what he <em>wouldn&#8217;t</em> do. He doesn&#8217;t generate a buy or sell recommendation every day just because he ran. He only speaks when there&#8217;s a real signal. A technical extreme. A material news event. A position that has drifted away from the thesis I wrote when I bought it. A banker who calls every day to say &#8220;nothing happened&#8221; is noise. One who calls only when something matters is signal.</p><p>I wrote that rule into Peter&#8217;s identity file, the document that gets loaded every time he wakes up. It is one line. It changes everything about how he behaves.</p><div><hr></div><h2><strong>What I didn&#8217;t expect</strong></h2><p><strong>The hardest part wasn&#8217;t the technology.</strong></p><p>It was decomposition. Breaking a vague intent like &#8220;I want to invest better&#8221; or &#8220;I want to write more regularly&#8221; into something an agent can actually execute. Not &#8220;do the analysis.&#8221; Instead: every morning at 8 AM, check these specific positions in these specific accounts, compare to these specific thresholds, and if a position is past a threshold, write a brief in this format.</p><p>The AI can execute almost anything you can describe precisely. The bottleneck is your ability to describe it. If you can&#8217;t say what &#8220;done&#8221; looks like in a way another human could verify, the agent can&#8217;t get you there either. That has turned out to be one of the most useful skills I&#8217;ve sharpened in the last six months, and it&#8217;s a product skill more than an engineering one.</p><p><strong>Memory architecture, not just memory.</strong></p><p>Giving an AI memory is not the same as giving it context. Memory is a pile of facts. Context is the right facts loaded at the right moment in the right shape.</p><p>My agents read from three layers. There&#8217;s a daily log that captures raw activity. There&#8217;s a long-term memory file that holds distilled lessons and the things about me that don&#8217;t change month to month. And there are project files (investment theses, weekly goals, identity documents) that get loaded only when they&#8217;re relevant to the task at hand. Nattu doesn&#8217;t see Peter&#8217;s investment theses. Peter doesn&#8217;t see my calendar. Each agent only loads what he needs to do his job.</p><p>This was the single biggest change. Before I separated memory from context, every agent session felt like talking to someone who had read everything once and forgotten most of it. After, each agent felt like someone who had been doing this job for me for months.</p><p><strong>Personality and constraint are part of the product.</strong></p><p>This is the thing I want to be clearest about, because most posts about AI agents skip it.</p><p>It is not enough to spin up a tool and ask it to &#8220;help with my calendar.&#8221; You have to tell the agent who he is. What he cares about. What he refuses to do. When he should stay quiet. What tone he speaks to me in. What I&#8217;m trying to accomplish this quarter, this week, today.</p><p>The agents that work for me work because each one has a written identity. Nattu has a Chief of Staff personality. He is direct, he respects my time, he doesn&#8217;t ask three follow-up questions when one will do, and he stays quiet when he has nothing useful to add. Peter has a different personality. Careful, conservative, slow to act, fast to flag. Each of them has a set of constraints written down that I revise every few weeks as I learn what&#8217;s actually useful.</p><p>The tool is the easy part. The product decisions (who is this agent, what does he do, what does he refuse to do, what does &#8220;done&#8221; look like) are the work. Those are the same decisions any good product manager makes about a feature. They&#8217;re just being made about a teammate now.</p><div><hr></div><h2><strong>What this is actually about</strong></h2><p>The tools have been good enough for a while. The bottleneck has shifted.</p><p>If you&#8217;re a product builder, an operator, or a domain expert, and you haven&#8217;t started building with these tools yet, I think the barrier is lower than you expect. The technical floor has dropped. The skill that matters is the same one that makes someone good at building any product: knowing what you want, knowing who it&#8217;s for, defining what done looks like, and being honest about what should happen when it fails.</p><p>I&#8217;m writing weekly about what I&#8217;m building. If you&#8217;re building like this, or thinking about it, I&#8217;d like to hear from you. Reply to this post or find me on X at @srinatar.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://blog.rememberloop.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Compounding Decisions! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[How I Finally Automated My Tesla Charging (With an AI Partner)]]></title><description><![CDATA[A story about Home Assistant frustration, NEM 3.0 math, three orphan crypto keys, and what it actually feels like to build something with AI]]></description><link>https://blog.rememberloop.com/p/how-i-finally-automated-my-tesla</link><guid isPermaLink="false">https://blog.rememberloop.com/p/how-i-finally-automated-my-tesla</guid><dc:creator><![CDATA[Sriram Natarajan]]></dc:creator><pubDate>Sun, 24 May 2026 04:42:32 GMT</pubDate><content:encoded><![CDATA[<p>I&#8217;ve had a Tesla Model Y and a Powerwall for over a year. I knew I was probably leaving money on the table with my charging schedule. PG&amp;E&#8217;s peak rates hit $0.38/kWh between 4&#8211;9 PM. My solar exports back at $0.08/kWh under NEM 3.0. The math is brutal &#8212; every kWh I self-consume is worth 4.75x more than one I sell back to the grid.</p><p>I knew what I wanted: a system that would charge the car at the right time, based on solar availability, grid prices, and battery state. Not a fixed schedule. An actual smart system.</p><p>I just couldn&#8217;t build it.</p><div><hr></div><h2>The Home Assistant Rabbit Hole</h2><p>Everyone in the Tesla/solar community points to Home Assistant. It&#8217;s open source, wildly flexible, has integrations for everything. On paper, it&#8217;s exactly what I needed.</p><p>In practice, it nearly broke me.</p><p>The documentation is sparse. Not &#8220;sparse&#8221; in the way that means you have to read carefully &#8212; sparse in the way that means critical steps are buried in three-year-old forum threads, half the integrations are community-maintained and outdated, and the error messages tell you nothing useful. I spent hours trying to get the Tesla integration working. I&#8217;d get partway through a setup, hit an undocumented wall, find a workaround on Reddit, try it, get a different error.</p><p>The frustrating part wasn&#8217;t that it was hard. It was that I couldn&#8217;t tell if I was one step away from it working or fundamentally doing it wrong. There was no one to ask. The docs assumed knowledge I didn&#8217;t have, and the community answers were scattered across years of forum posts that may or may not still apply.</p><p>I gave up. Not because I&#8217;m not technical enough &#8212; I&#8217;ve been hacking with technology for years. I gave up because the feedback loop was broken and I had no co-pilot.</p><div><hr></div><h2>Starting Over, Differently</h2><p>A few months later, I came back to the problem &#8212; this time with Nattu (my AI assistant running on OpenClaw) as a full collaborator.</p><p>The difference was immediate. Instead of hunting through docs alone, I could say: &#8220;Here&#8217;s what I&#8217;m trying to do. Here&#8217;s what I know about the Tesla API. Help me figure out the right approach.&#8221; And get back not just an answer, but a reasoned one with trade-offs explained.</p><p>I came in with a wish list a mile long: never charge during peak, prefer solar window on weekday afternoons, charge overnight on weekdays at the off-peak rate, opportunistically grab solar on weekends, maintain a battery floor, handle VPP dispatch events, defer to manual overrides. Every rule felt obvious in isolation.</p><p>I built all of them.</p><p>That was the first mistake.</p><div><hr></div><h2>Building the Actual System</h2><p>The stack we landed on:</p><ul><li><p><strong>Tesla Fleet API</strong> for reads &#8212; battery %, solar production, home load, Powerwall state, grid flow</p></li><li><p><strong>tesla-control binary</strong> for signed vehicle commands (start, stop, set amps)</p></li><li><p><strong>A virtual key paired to the car</strong> via my own domain (ev.srinatar.xyz) &#8212; Tesla&#8217;s new required auth model for any third-party command</p></li><li><p><code>charger.py</code> &#8212; Python script holding the decision logic</p></li><li><p><strong>launchd</strong> running it every 30 minutes as a background service on my Mac mini</p></li></ul><p>Let me tell you about the virtual key.</p><p>Tesla&#8217;s newer API requires you to generate an ECDSA keypair, host the public key at a very specific path (<code>/.well-known/appspecific/com.tesla.3p.public-key.pem</code>) on a domain you control, register that domain with Tesla via their partner API, and then pair the key to the car from the Tesla app via a specially-crafted URL. The car then asks you to approve the pairing on the touchscreen. Once paired, every vehicle command has to be cryptographically signed by your private key before Tesla will accept it.</p><p>This is a <em>good</em> security model. It is also a security model that gives you many opportunities to footgun yourself.</p><p>Here&#8217;s where I ended up at one point: three different public keys on disk, no idea which one was paired with the car, an nginx serving one of them publicly but no matching private key anywhere on the system, and a Cloudflare tunnel that &#8212; I would later discover &#8212; didn&#8217;t even have a public hostname route configured, so the public URL where Tesla expected to fetch my key was timing out from the outside world.</p><p>Tesla&#8217;s documentation does not warn you about any of this.</p><p>Nattu helped me untangle it: ran a hash on every <code>.pem</code> file in my home directory, cross-referenced which key matched the one nginx was serving, figured out the Cloudflare tunnel was misconfigured, walked me through generating a fresh keypair from scratch, hosting it, registering it with Tesla, and pairing it. End-to-end signed command working: <code>tesla-control charging-set-amps 14</code> &#8594; car responds, amps change, exit code 0.</p><p>That was the high point of the project &#8212; not because automated charging is technically impressive, but because the kind of debugging it required (a six-step distributed system spanning my Mac, nginx, Cloudflare, Tesla&#8217;s servers, and the car itself) is exactly the kind of thing I would have given up on, alone, like I gave up on Home Assistant.</p><div><hr></div><h2>The Less-Is-More Lesson</h2><p>While the auth was being wrangled, the decision logic was running. And it kept doing things I didn&#8217;t want.</p><p>I&#8217;d manually start a charge at 2 PM because I knew solar was abundant. Thirty minutes later, <code>charger.py</code> would stop it because some &#8220;insufficient solar/PW headroom&#8221; rule fired. I&#8217;d start it again. It would stop it again. The car got fewer kWh than if the system hadn&#8217;t existed at all.</p><p>The smart system was actively making my life worse.</p><p>The fix was uncomfortable: we ripped out almost every rule. The current <code>charger.py</code> does exactly one proactive thing &#8212; it stops charging during peak hours (4&#8211;9 PM). All other times, it stays out of the way. If I want to charge, I charge. If I&#8217;m running a special solar-only week before a road trip, that&#8217;s a separate explicit override. When the peak-hour stop fires, it sends me a Telegram notification so I&#8217;m never surprised.</p><p>The lesson was clear and slightly humbling: the bottleneck wasn&#8217;t smarter logic. The bottleneck was a clean line between &#8220;what the system decides&#8221; and &#8220;what I decide.&#8221; Once those stopped fighting each other, everything worked.</p><p>The 30-minute launchd job still fires. Most of the time it logs the state, looks around, and does nothing. That&#8217;s the system working correctly.</p><div><hr></div><h2>What Building With AI Actually Felt Like</h2><p>This is the part I want you to take-away from this post. </p><p>I&#8217;ve used AI tools for a long time. I use them to write, to research, to summarize. That&#8217;s using AI. This was different &#8212; this was <em>building with</em> AI.</p><p>The difference is that Nattu held context across the entire project. When I hit a problem with the Fleet API auth, I didn&#8217;t have to re-explain the whole setup. When the over-engineered first version was misbehaving, Nattu remembered the NEM 3.0 math and pushed back when I tried to bolt on more rules instead of removing them. When I got frustrated chasing the three-orphan-keys mystery and wanted to take a shortcut, I got an honest read on why that shortcut would leave me worse off.</p><p>It felt less like querying a tool and more like working with someone who was actually invested in getting it right.</p><p>The Home Assistant failure wasn&#8217;t really about the software. It was about trying to navigate complexity alone, without a feedback loop, without someone to reason through it with. That&#8217;s what was missing.</p><div><hr></div><h2>What&#8217;s Next</h2><p>The simple system works. Now that signed vehicle commands are unblocked, the next step is the version I actually wanted from the start: <strong>adaptive amp control</strong>. Read real-time solar output and home load every 30 minutes. Compute the surplus. Set the car&#8217;s charge amps to match &#8212; pause if surplus is tiny, ramp up to 18A when solar is pouring in. Manage the Powerwall so it doesn&#8217;t hit 100% before peak hours, preserving headroom to absorb the late-afternoon solar that would otherwise export at $0.08/kWh.</p><p>The economics are simple: turning a $0.08 export into a $0.36 self-consumed kWh, every kWh, every day the sun shines. Over a year, that&#8217;s real money.</p><p>But this time I&#8217;m building it with the lesson freshly learned: one rule at a time, each one earning its keep, with a kill switch in plain sight.</p><p>The deeper lesson I keep coming back to: the bottleneck to building useful things with AI isn&#8217;t intelligence. It&#8217;s the quality of collaboration. The projects where I&#8217;ve made real progress are the ones where I brought a real problem, stayed in the conversation, pushed back when something didn&#8217;t feel right, and was willing to delete code that wasn&#8217;t earning its complexity.</p><p>That&#8217;s not a new skill. That&#8217;s just how good work gets done &#8212; with a partner who&#8217;s paying attention.</p><div><hr></div><p><em>If you&#8217;re trying to do something similar with Tesla + solar, I&#8217;m happy to share the scripts. Find me on X at @srinatar.</em></p>]]></content:encoded></item><item><title><![CDATA[How I Optimized Charging My Tesla on Sunshine]]></title><description><![CDATA[Multiple attempts of frustration, then a single afternoon with my AI agent.]]></description><link>https://blog.rememberloop.com/p/how-i-optimized-charging-my-tesla</link><guid isPermaLink="false">https://blog.rememberloop.com/p/how-i-optimized-charging-my-tesla</guid><dc:creator><![CDATA[Sriram Natarajan]]></dc:creator><pubDate>Sun, 17 May 2026 06:15:41 GMT</pubDate><content:encoded><![CDATA[<p>Recently, I came back from a five-day trip. While I was gone, my Tesla charged itself off the Sun every afternoon and didn&#8217;t pull from the grid at night. I never opened the app. My power bill for the week was lower than a normal week at home.</p><p>It took me multiple months and three attempts to get to this point. I almost gave up for good. The third attempt worked, and it worked because I stopped trying to do it alone.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://blog.rememberloop.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Compounding Decisions! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>Here&#8217;s what is different now.</p><h3>The two times I gave up</h3><p>The setup I wanted is simple. When the car is home all day, charge it from solar, not the grid, and not overnight. Keep the Powerwall full enough to run the house through the night. This combination maximizes my solar investment and lowers my grid usage.</p><p>The first time I tried, I went the Home Assistant route. Everyone in the Tesla and solar communities swears by it. In practice, the documentation was scattered across three years of forum posts, and half the integrations were community-maintained and out of date. I&#8217;d get partway through a setup, hit a wall, find a workaround on Reddit or in the Home Assistant community forum, try it, get a different error. I couldn&#8217;t tell if I was one step from working or fundamentally on the wrong path. I closed the tab after a couple of weeks into it.</p><p>The second time I wrote my own scripts against the Tesla API. Reads were easy. Writes were not. Tesla had rolled out an auth model where any command that actually does something to the car has to be cryptographically signed by a virtual key paired to the vehicle. Generate the keypair, host the public key on a domain you control, register the domain with Tesla, walk the pairing flow through the app, approve it on the touchscreen. The documentation existed. It had quietly load-bearing gaps. I gave up again.</p><h3>The third attempt</h3><p>The third attempt happened because I sat down with Nattu and treated him as a co-builder, not a search engine.</p><p>Who is Nattu? Nattu is my AI agent. He runs on OpenClaw, a multi-agent gateway I&#8217;ve been hacking on. For this post, what matters is that he held context across the whole project and pushed back when I tried to overcomplicate things.</p><p>The auth nightmare took an afternoon. We found three different public keys from my previous attempts on my Mac. None of them had a matching private key anywhere on disk. The Cloudflare tunnel that was supposed to expose my Mac to Tesla&#8217;s servers had no public hostname route configured, so the URL Tesla was supposed to fetch was timing out from the outside. </p><p>We hashed every .pem file, figured out which was which, generated a fresh keypair, fixed the tunnel, re-registered with Tesla, re-paired. End-to-end, this would have taken me weeks alone, but we got it done in under 30 minutes.</p><h3>What the system actually does</h3><p>What the system actually does is small. </p><p>Most days, he stays out of the way. The only proactive rule is: stop charging during peak hours, 4 PM to 9 PM. Outside of that window, if I want to charge, I charge. If I don&#8217;t, nothing happens.</p><p>On Weekends or When I&#8217;m traveling, Nattu has access to my calendar and knows when I&#8217;m returning. He knows the rules change. For example, he doesn&#8217;t do overnight charging but optimizes to charge only from solar during the day (between 10 AM and 4 PM), while keeping the Powerwall above 80%. Once he knows I&#8217;m back, he automatically reverts to the normal schedule.</p><p>Every decision Nattu makes regarding charging, he sends me a message on Telegram. &#8220;Solar surplus 4.2 kW, PW at 84%, starting at 14A.&#8221; &#8220;PW dropped to 78%, pausing.&#8221; I don&#8217;t read them live. They sit there as a log I can scroll through later.</p><h3>The trip and the bug</h3><p>A few weeks into having travel mode running, I had to fly out for a customer visit on short notice. Nattu saw the trip on my calendar and switched to travel mode that night.</p><p>The trip went well. The car charged from the sun, Powerwall stayed full, the bill came in lower than a normal week. I got home Tuesday night to a full battery.</p><p>Wednesday morning, the day everything was supposed to restore, I plugged the car in around 11 AM. At 11:40 the car stopped charging. I started it from the app. At 2:40 it stopped again.</p><p>Three minutes of reading logs told me the whole story. The travel script had auto-restored the normal charger correctly. It just hadn&#8217;t unloaded itself. Two scripts were running every 30 minutes, one that allowed charging, the other still enforcing the travel-mode window for a trip that had ended the day before. They kept stopping my charge.</p><p>I had built the travel mode and the restore. I&#8217;d forgotten to build the cleanup that removes the travel script from the schedule when the trip ends.</p><p>The fix was small. Unload the old job, verify it actually unloaded, alert me on Telegram if it didn&#8217;t. We added a rule to our shared notes: when you build something that hot-swaps logic, write the cleanup before you write the activation.</p><p>The Telegram log mattered here. Without it, I&#8217;d have caught the bug weeks later, squinting at a higher-than-usual power bill. With it, I caught it in three minutes by scrolling back.</p><h3>Stepping back</h3><p>Two things stayed with me from this.</p><p>Working with an agent is different from using AI to write or research. I&#8217;ve done a lot of that. This was different because Nattu held context across weeks of stop-and-start work. He pushed back when I tried to bolt on more rules instead of removing them. He kept me honest about whether the system was actually working or just looked like it was. The Home Assistant attempt failed because I had no one to think with. This one worked because I did.</p><p>The other thing is the list. The Tesla project sat on my list for months. I have a long list of problems like this one. Each one is solvable in theory. None of them fit into the hours I actually have. </p><p>The point of this post isn&#8217;t that you should automate your Tesla charging with an AI agent. It&#8217;s that the problems on your list that feel just out of reach probably aren&#8217;t, anymore. </p><p>Sit down with an AI agent as a partner. Start with the smallest version of what you want, and only add a rule once you&#8217;ve felt the absence of it.</p><p>Build the thing. Trust it. Verify it. In that order.</p><p><em><strong>If you&#8217;re trying something similar, find me on X at <a href="https://x.com/srinatar">@srinatar</a>.</strong></em></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://blog.rememberloop.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Compounding Decisions! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item></channel></rss>