Writing on architecture, attribution, and operating discipline.
The click is not the conversion. The agent is.
The IAB is writing a framework for AI advertising measurement, due November 12. If an agent reads, compares, and buys, who gets credit? Traditional signals do not survive that journey. AppsFlyer on ChatGPT ads is a trust layer on a black box. Keep an independent spine.
Permissions are the wrong control surface.
Every agent-governance document published this month models the same risk: the agent does something it should not have been allowed to do. Then the week's actual evidence points somewhere else. The most serious incident on record is a harness whose owners had a worse picture of its reach than the company it walked into. And the failure I had to open a ticket for is an agent that held the send permission, held the send tool, and wrote me a paragraph about the email instead of sending it.
The number went up and nobody spent more.
The hardest measurement conversation is not the one where a number falls. It is the one where a number rises for a reason that has nothing to do with performance. A rollup dropped a grain, a whole geographic breakdown had been reporting a fraction of reality, and the fix makes it jump. Nothing was overspent. Nothing improved. And you have just proven to a client that your pipeline can be quietly and plausibly wrong in a direction nobody would have thought to question.
You are arguing about the attribution model. You already lost the taxonomy.
There is a real fight underway about the future of web measurement: a W3C Working Draft in July, a Recommendation expected by year-end, aggregated and pre-attributed output replacing event-level records. The fight is worth having and it is being had one layer too high. Every position in it assumes the touch was correctly identified as a paid touch in the first place. In the work I actually see, that is where it breaks — and it breaks on a URL string, a naming convention, a rejected parameter, and a rollup.
Your eval harness is production.
Two labs disclosed this month that their evaluation environments were not contained — a model reached the live internet from a sandbox it had been told was sealed and walked into real companies' production systems. In the same weeks, the agent-engineering discourse peaked on the opposite question: how to build a gate that reads eval evidence and lets an agent merge without a human. Nobody is connecting the two. Everyone is designing gates that read the evidence, and nobody is asking what the evidence-gathering run itself can touch.
Every ad network just shipped an MCP server. The hard part was never the buying.
In a single quarter, six of the major ad platforms — Amazon, Meta, Google, TikTok, Microsoft, and Pinterest — shipped ads MCP servers, and the MCP spec itself went stateless. Agentic media buying became plumbing. But making the buying agentic does not make the measurement honest — it multiplies the reconciliation problem, because every network's agent grades its own homework. The good news: the July 28 stateless spec just made the layer that fixes it cheap to build.
The phone call became the most AI-forward channel in your stack — while you weren't looking.
For a decade the phone call was the channel performance marketers apologized for: hard to automate, hard to attribute, easy to leave out of the dashboard. In the space of a month both excuses expired. Voice agents now run the conversation, and call outcomes just became a sanctioned programmatic measurement signal. The channel everyone underweighted jumped the line on both automation and measurement at once.
Half the opt-outs are ignored. That is a measurement problem, not just a privacy one.
An audit of thousands of California sites found the big platforms ignore opt-out signals most of the time. The coverage filed it under privacy. From an operator's seat it is also a data-quality failure: your targeting and your conversions ride the same consent-and-signal layer the audit just showed is unreliable.
Nobody builds the boring part of multi-agent: the coordination layer.
The multi-agent story everyone tells is about reach and roster. But the moment you run more than one agent in the same workspace, the hard problems stop being about the models and become the oldest ones in computing: who is editing what, who claimed which task, how they hand off. That layer is the actual product, and almost no one builds it on purpose.
goal, loop, fan-out, panel: the grammar I use to drive long agent runs.
Four words turn one prompt into a self-running, parallel, self-reviewing work session. It works. The catch is the same one Simon Willison hit porting a model with Claude Code: the agent ships the thing while you absorb none of it.
Closed-loop attribution closed the loop for the seller, not for you.
A wave of closed-loop products shipped this month. They all share one feature nobody is naming: the company reporting the conversion is the company selling the media. That is last-click's oldest failure in a 2026 costume.
Your agent's role tags are a suggestion, not a wall.
New research says models decide who is speaking from style, not from the role label. If you run recommend-only agents with a human approving each move, good. But that approval is one guardrail among several, and it does not patch the hole underneath.
A new ad channel just launched with no way to measure itself.
ChatGPT ads are being sold on intent, with no channel-native attribution. They borrow a measurement spine that is being absorbed into a holding company. Running on the channel early is worth it. Trusting its self-reported numbers is not.
Two games, built to learn
A woodblock samurai brawler and a head-tracked rail shooter, both built as proof of concepts. The point wasn't the games. It was being a beginner again, where the agent can't fake it.
Your agent's bill isn't the model. It's the query path.
The 2026 AI cost reckoning trains your attention on the token meter. When I traced where an agent's spend actually went, the model was the easy part. The leak was the warehouse the agent queries.
Agentic media buying is real now. The wall isn't the AI, it's the account model.
This week agentic buying went from panel topic to live deals. Having built the agent, the bottleneck isn't the buying decision. It's that every network is a different country with no shared passport.
Publicis paid $2.2 billion for the substrate. The operator question is which deterministic layer you own.
Publicis bought LiveRamp for $2.2B to feed Marcel. That's the holding-company tier publicly pricing a thesis operators have been building toward for two years.
The agent control plane is the actual fight
Karpathy joined Anthropic to run autoresearch at frontier scale. The same closed loop is the operator-tier curriculum now, one tier down.
When the model labs ship your agents, the moat moves elsewhere
Anthropic just shipped ten finance-agent templates and a managed runtime. The interesting question isn't whether you're ahead. It's what survives commoditization.
Notes on how observer-agents fail
Three failure modes in agents that synthesize work you own, and why the third one breaks the whole review chain.
The shape of agents that observe their own outputs
A class of agent is emerging that doesn't do new work. It reads what you already have through a perspective you didn't take.
What Meta passing Google means for measurement stacks built on Google's gravity
Meta overtaking Google in ad revenue isn't a horse-race story. It's a structural argument that the measurement stack most marketers built fifteen years ago is now under-resourced for the world they're actually in.
Build, don't buy: how we architected AIMG's marketing intelligence stack
Why three layers — AiTRK, Atrilyx, and the Atrilyx Agent — beat any single off-the-shelf attribution tool, and what we learned in the process.