Two months of building an AI support agent in public: what actually changed my mind
Sixty days into documenting this build in public, the lessons split cleanly into two piles: what changed about how we build the agent, and what changed about how we talk about building it. The second pile surprised us more, and it maps almost exactly onto owned, shared, and earned media — the three channels every piece of this series actually traveled through.
We didn't set out to run this as a deliberate content strategy when it started. It became one gradually, once it became obvious that the writing itself was surfacing problems and questions we wouldn't have found just by building quietly.
Owned
The blog and the long-form breakdowns are the owned channel — the place we control completely, where the record actually lives. Four things earned a permanent place there. First: the model was never the actual bottleneck; tool design, state tracking, and policy consistently mattered more than which model sat behind the conversation. Second: "autonomous" is a spectrum, not a switch, and customers can feel which point on it they're talking to. Third: the unglamorous parts — notification icons, embed scripts, launcher hover states — decide trust before the AI ever gets a chance to prove itself. Fourth: the best evaluation was never "is it smart," it was "where exactly does it fail, and does it fail safely," a harder question but the only one that predicts what happens once real customers, not test conversations, start using it.
Writing these up properly, at the length they actually deserved rather than the length that was quick to publish, forced a kind of rigor the daily posts never demanded. You can gesture at a policy engine in one line on a fast-moving feed. You can't gesture at it in a piece meant to sit on the site for months, and drafting the longer version repeatedly surfaced gaps in our own reasoning that the short version had let us skip past.
Shared
The daily posts were the shared channel — shorter, faster, meant to travel on their own rather than sit in an archive. What we learned there was less about the content and more about the format: a single concrete claim, stated plainly, travels further than a hedge wrapped in three qualifiers. The posts that got the most genuine engagement weren't the ones announcing progress, they were the ones admitting a specific failure mode by name — a timeout, a contradiction the agent almost missed, a security gap we closed before anyone hit it. People share specificity. They scroll past reassurance.
We also learned, less comfortably, that the shared channel rewards a kind of confidence the owned channel shouldn't. A post that states a strong opinion plainly outperforms one that states the same opinion with appropriate nuance — and the discipline that took real practice wasn't writing the confident version, it was making sure the confident version stayed honest, with the nuance moved to the long-form piece rather than dropped entirely.
Earned
The earned channel was the smallest and the most valuable — replies, questions, the occasional "we hit the same wall" from someone building something adjacent. None of it was requested, and none of it was predictable in advance, which is exactly what makes it earned rather than owned or shared. It's also the channel that corrected us fastest: a question about a specific failure mode we hadn't written up yet is a more useful signal than any metric we track internally, because it tells us what an outside reader actually needed to know and didn't yet.
More than one of the long-form pieces in this series exists specifically because a reply on a shorter post asked a question we didn't have a clean answer for yet. That's the earned channel doing something the other two can't: telling you, in real time, exactly where the gap in your own explanation actually is, from someone with no stake in you looking good.
What ties it together
A specific example of each channel doing its job
The clearest owned-channel example was the piece on policy engines earlier in this series. It started as a single line in a daily post — a claim that model choice barely moved resolution rate once a policy layer existed — and the long-form version only came together after we forced ourselves to actually produce the before-and-after numbers behind that claim, rather than relying on the general sense that it was true. Writing the owned version is what turned an intuition into something we'd trust enough to publish with actual measurements attached.
The clearest shared-channel example was the opposite kind of lesson. A short post admitting that an earlier version of our escalation handoff was actively making the human's job harder — arriving with none of the context the agent had already gathered — outperformed every purely positive post we'd written that month, by a wide margin, on every engagement measure we track. We hadn't planned it as a strategy; it was just an honest admission that happened to travel, and the pattern held up enough afterward that we started deliberately looking for the next honest admission worth making, rather than treating it as a one-off.
The clearest earned-channel example was a reply on that same post, from someone building a completely unrelated kind of agent, asking a pointed question about how we actually detect a bad handoff after the fact rather than just how we designed the good one. We didn't have a clean answer at the time. The piece on the anatomy of support tickets earlier in this series exists partly because answering that one reply properly turned out to require writing something much longer than a reply.
None of these three examples were planned as a system when the series started. What we noticed, after enough of them accumulated, is that owned, shared, and earned weren't really three separate content channels doing three separate jobs — they were three separate feedback loops, each catching a different kind of gap in our own understanding, and the writing itself turned out to be doing almost as much work as any single engineering sprint in the same period.
Common questions
Was building in public ever a liability — did competitors or bad actors use any of this against you? Less than we expected going in. The specific failure modes we've written about are exactly the kind of thing a serious competitor already assumes exists in any agent at this stage — writing about them publicly cost us less in competitive exposure than it gained us in credibility with buyers who've been burned by vendors that pretend these problems don't exist.
How do you decide what not to publish? Anything that would expose a specific customer's data or a genuinely unpatched vulnerability doesn't get written up until it's fixed, without exception. Short of that line, our bias has shifted toward publishing more than we started with, because the earned-channel value of an honest failure post has consistently outweighed the discomfort of admitting the failure.
Does this series actually influence the roadmap, or is it a separate track from the real engineering decisions? It's influenced the roadmap directly more than once, most visibly through the earned channel — a specific question from a reader is what prompted the deeper dive into escalation-handoff quality that became its own piece. We wouldn't call it a separate track; if anything, writing the long-form pieces has become one of the more reliable ways we find gaps that a normal engineering retro doesn't surface, because writing for an outside reader forces a level of explicit reasoning internal shorthand usually skips.
What would you tell someone starting a similar build-in-public series today? Don't wait for the content plan to feel complete before starting, and expect the shared channel to teach you more about your own writing than about your own product for the first few weeks. The value compounds once the three channels start feeding each other — a shared post surfaces a question, the question becomes a long-form piece, the long-form piece gets shared again — but that loop takes a few cycles to start turning on its own.
The three things we'd do differently next time
Two months in, three specific things stand out as changes we'd make if we were starting this series over, and none of them are about topics we should have covered — they're about how we approached the writing itself.
The first is that we'd instrument the shared channel from day one the same way we eventually instrumented the widget's customer journey — tracking which specific claims travel and which don't, rather than watching engagement in a general, undifferentiated way. We didn't start doing this until several weeks in, and looking back at the earlier posts, we're fairly confident we could have identified the honest-admission pattern — that specific failures travel further than general progress updates — several weeks sooner than we actually did, simply by treating our own content the way we treat any other product surface we measure.
The second is that we'd separate the owned and shared versions of a claim more deliberately from the start, instead of letting the shared version sometimes be the only version that existed for weeks before the long-form piece caught up. More than one daily post made a claim confidently that the corresponding long-form piece later had to walk back slightly or add real nuance to, once we'd actually done the work to verify it properly. That's a fine outcome when it happens occasionally and gets corrected quickly, but doing it deliberately — publishing the confident short version and treating the long-form version as "eventually" rather than "soon" — created a gap where the fast-traveling claim was ahead of the verified one for longer than it should have been.
The third is that we'd engage more actively with the earned channel earlier, rather than treating replies and questions as a pleasant side effect rather than a primary source of direction. Several of the long-form pieces in this series exist because a reply eventually forced the question, but looking back, similar questions showed up in earlier replies that we didn't follow up on with the same seriousness, simply because we hadn't yet recognized that channel as a legitimate source of what to write next rather than just a nice signal that people were reading.
None of these three are dramatic reversals — we're not restructuring the series, and the core thesis hasn't changed. They're the kind of adjustment that only becomes visible in hindsight, once enough of the pattern has accumulated to see it clearly, which is itself a small instance of the larger point this whole piece has been making: writing the thing down, consistently, over time, surfaces gaps that thinking about it privately never quite does.
If there's one adjustment we'd recommend to anyone starting something similar today, informed by getting this wrong ourselves for the first several weeks, it's this: build the measurement and the habit of following up on outside questions into the process from the very first post, rather than waiting until the pattern is obvious enough to notice by accident. It was obvious to us eventually. It would have been more useful several weeks earlier.
None of this changes the underlying thesis, which the whole series has been circling from the start: an agent's job isn't to sound helpful, it's to change the state of the world for the customer, safely, and be able to show its work afterward. What changed is how we think about explaining that thesis — the owned channel for the version that has to hold up over time, the shared channel for the version that has to travel, and the earned channel for finding out, honestly, where either version still doesn't hold up.
Month three starts from here — deeper into policy engines, real evaluation numbers, and what happens the first time the agent handles a genuinely angry customer at scale. If there's a specific failure mode you want broken down next, that's the earned channel doing its job — say so.
We'll keep the same discipline going into month three that got us here: publish the owned version once we've actually verified it, not just once it sounds right; let the shared version travel on its own without dressing it up past what we're confident in; and treat every earned reply as a real input into what gets written next, not a metric to glance at and move past.