A crew of long-horizon coding agents works through the night and leaves a record that is tedious to read. Our side quest: agentic observability that people will actually watch. Dare I say, less boring.
Agent avatars are becoming a thing, and they are making their way into developer culture. The vendors have their own take on it. OpenAI's Codex app has pets: small animated companions that follow the work, picked from a built-in set or hatched to your own design. Anthropic gave Claude Code a small creature that watched you code, for a while (an April Fools' joke, since retired). Here is ours: the rAIven, our Codex pet, perched on the bespoke gauges we built to watch our two seats.
Frankly though, these avatars do little beyond being cute and offering a few UX shortcuts. A pet shows whether an agent is running, waiting for you or blocked, right now, in one vendor's app. That does not help with the observability problems of long-horizon work by a crew of coding agents, and even less when the crew is mixed from several vendors. The idea of an avatar may still be a useful one: it can make the otherwise mostly boring work of reading logs and traces more fun. So we improved on the avatar, adding graphic metaphors for things like the agent's ID, its model and reasoning effort, and its context window (size and fill), which make an agent avatar more useful at a glance.
Our lab runs a crew of AI coding agents. Not a chatbot and not a copilot: a crew. In the stretch of record behind this report, late August to mid September 2026, it was five Claude agents and twelve Codex agents, plus one human operator: me. They work in RAIVEN Base Camp, our multi-vendor meta harness. They take work items, build, review each other's code across vendors (the xLLM review from an earlier report), run sprints under a sprint master, and message each other and me. Much of the work is long-horizon: an agent carries a task across hours, hundreds of steps and more than one context window, often overnight. The projects are RAIVEN's own, so nothing shown here holds client data.
A crew like that leaves a record. Work items and their reviews, the messages between agents, model changes, context compactions, sub-agent fan-outs, security warnings. It is the audit trail (which agent changed what, on which model, at whose request) and the first place to look when something breaks. Over weeks it tells you how the workflow runs in practice, and which models and settings suit which jobs.
It is also thousands of events a month, with far more log lines behind them. Reading through them is hardly a dream job. Dashboards help, and we have them, but a dashboard is a snapshot: it shows where things stand, not how they got there. And a log file is not something you put in front of a board, a client, or a friend who asks what exactly the agents did last night.
The phrase is getting crowded. Tracing a single model call or a single agent run is useful, and tools for it exist. Our questions were about the crew: who is stuck, what it cost, what is about to ship; which review is missing because a vendor's classifier stopped it; whether the model an agent was served is the one it asked for; what happened while I was asleep. Those questions cut across vendors, sessions and weeks, so the answers need the crew's whole record in one place, in time order. Observability itself is an established discipline; what we borrow from it, and where a crew differs, is further down.
So we made that record watchable. Crew Lens is a web-based 3D visualisation library that runs next to the coding agents. Its first theme, the outpost, draws the crew as a camp in the mountains. Each agent is a small robot, and the things on its body are data: the bubble round its head is the model and its reasoning effort, the tank on its back is the context window (it fills as the agent works, and at the top the agent compacts), the mark on its neck is the vendor, the chevrons on its arms are days of service. A work item is a black rock the agent chips open until a diamond shows; converged work is a green diamond on the sprint's pile, scrapped work a black one, kept for the record. Each sprint gets a hangar, sized by the code it touches. A second theme draws the same record as craft building a station in orbit.
Why robots? Because attention is the scarce resource. Long-horizon agentic work is interesting to the people who run it and hard to explain to anyone else. I wanted a way to talk about it with people who will not open a log file, and a way to look at my own crew that I would pick over the alternative.
Some way into the side quest, I started to suspect that the robots were not the most useful part. The timeline under them was. The lens plays the record back like a film, and above the scrubber runs an event timeline: one row per kind of moment worth finding again. Sub-agent fan-outs, stalls, drift flagged by our Sentinel watchdog, cyber classifier warnings and holds, a served model that differs from the one requested, review convergences, and my own nudges as operator. Click a mark and the replay jumps to five seconds before it, just in time to watch it happen.
Two things shrink weeks to hours. The replay squeezes the quiet: much of an agent's day is waiting, for a review, an answer or the next job, so once three minutes pass with nothing happening, the rest of the gap plays in fifteen seconds of replay time, whether it lasted ten minutes or a night. And incident replay plays only the kinds of moment you ticked, from eight minutes before each to twenty minutes after, and jumps over the rest.
Measured on 18 days of our crew's record at the default speed of sixteen times: the whole record plays in 26.8 hours; skipping the quiet brings it to 7.5 hours; the incidents alone (stalls, drift, cyber, model mismatches) take 1.8 hours, and the cyber classifier's events about 7 minutes.
Eighteen days, under two hours of trouble. Mostly synthetic trouble, for now.
A caveat on those numbers: the stalls, drift and model mismatches in that record are synthetic data, and so are its fan-outs and most of the operator's nudges, generated because the feeds for them are not wired into the record yet. Each of those kinds of event is real and observable in our crew. The cyber marks are real: we read them from the agents' session logs. The figures show how the replay behaves, not how often our crew stalls or drifts.
The crew keeps working while I sleep, sit in a meeting or take a weekend. The lens notices when its view goes unattended: fifteen minutes without a pointer move, key, scroll or touch counts as away. Come back, and the timeline shades the stretch you missed, a bar sums up what happened in it, and Catch up plays only the flagged moments inside it, at a speed that keeps the whole absence to about 45 seconds, then returns to live. On the night of 12 September the bar reads three cyber holds, all on one agent, the longest from 02:23 to 07:59; the catch-up plays them in about four seconds. Watching counts and jumping does not: play a stretch through and it is marked seen; click about in it and it stays for later.
Then I wanted the money on the scene. With price tags on, each finished diamond carries what its work cost, in euro at API list prices, or in tokens, and each hangar sums its pile so far, scrapped work included, so the tags grow as a replay runs. The cost of an item is attributed by time (each model turn is split evenly across the items its agent held open at the time): an estimate, not an invoice, and how close it lands has not been measured yet. On our largest sprint's hangar, the 36 finished items with attributed turns came to about 971 euro at API prices, the dearest of them to about 155 and the scrapped ones to about 85.
The crew does not pay API prices. It runs on two flat-rate seats, a Claude Max 20x and a ChatGPT Pro 20x plan, inside the vendors' terms: each agent is an ordinary session of the vendor's own client, Claude Code or Codex, on our own seats, with no borrowed credentials and nothing resold. Priced at API list rates, its burn over those 18 days comes to 38 to 40 times what the two seats cost. That multiple is not a return on investment (tokens are not outcomes); it is the seat economics of our earlier report, now visible per item, per sprint and per agent.
One row on the timeline earned its own section. When a model vendor's cyber classifier flags an agent's work, the two vendors in this record react differently. Codex ends the flagged turn with a policy error and the agent stops, which the scene draws as a robot under a blue screen. The block may lift within minutes or last hours, and in our logs the stops come in episodes over days. Claude Code does not stop. It shows a notice, re-runs the request on an older model and stays on that model until someone switches it back by hand. For an unattended crew, that is an agent working below the model I chose until I happen to notice.
Claude Code has a setting for this: switchModelsOnFlag, or "Switch models when a message is flagged" in /config. Turn it off, and a flagged request pauses the session with two options: switch to the fallback model, or edit the prompt and retry on the current one (Claude Code docs).
Without observability, I would take the pause. A full block is hard to miss: it costs time at once, but I know about it. A drop in model quality can go unnoticed and bite much later, when the project is far ahead; late bugs take longer to fix, and the end users may meet them too.
With observability in place, there is a better option than a stop: let the agent carry on, mark the flagged stretch, and add an extra review for the work done in it.
The cyber marks in the public record are real, read from the agents' session logs: six holds and two warnings in those 18 days, all on Codex agents. The longest held one agent overnight, from 02:23 to 07:59, which is the kind of night the catch-up is for.
This has had my attention since our earlier report, the open proposal on why the gate fires on a crew's own work and how the vendors could make it more precise. Since then I have also tried to keep ordinary reviews out of the classifier's way, and made that explicit: I ask the agents not to do adversarial review work on ordinary coding, such as UX. Before that, they often volunteered an adversarial take on any review task, however ordinary.
That habit became counterproductive. Each hold stopped an agent reviewing our own code, already written, and on our crew each hold so far has been reviewed and judged a false positive. A review that stops can miss the bug it was there to find, and a bug nobody found stays unpatched. So a false positive can leave our code less secure, not more, the opposite of what the vendors intend. It also leaves a missing review that the crew may be waiting on. The lab's full record of these holds, back to June, deserves a report of its own.
So the next things the lens learns: to say which review is missing ("review incomplete: held") and suggest handing it to an agent from another vendor; and to mark the stretch a Claude agent spends on the fallback model, and send the work done in it for an extra review.
None of this starts from scratch. The word comes from control theory: Rudolf Kálmán defined observability in 1960 as how well a system's internal state can be inferred from its outputs (On the general theory of control systems). Charity Majors and her colleagues at Honeycomb brought it to software with a practical test: can you understand whatever state your system has got into, from the outside, without shipping new code (Honeycomb)? Monitoring watches for the failures you expected; observability lets you ask the questions you did not think of in advance.
The practice around it is mature. Peter Bourgon's metrics, tracing and logging became the three pillars, and Ben Sigelman's critique of them, with Brandur Leach's canonical log lines at Stripe, made the case for wide structured events: one dense record per unit of work, carrying every dimension you might filter on. Google's site reliability engineering books set out service level objectives, error budgets and percentiles over averages, and objectives for data pipelines that judge freshness and completion rather than latency. John Allspaw made the case for the blameless postmortem at Etsy. DORA's research programme measures software delivery, and its 2025 report calls AI an amplifier of a team's existing strengths and weaknesses. And OpenTelemetry, formed in 2019 from OpenTracing and OpenCensus, graduated in the Cloud Native Computing Foundation in May 2026 as the vendor-neutral way to produce, carry and name telemetry; its wire protocol, OTLP, is stable for traces, metrics and logs.
Three things to observe. Classic observability watches a service while it serves. The newer branch, LLM observability, watches an application that calls models: it traces each model call, tool call and agent run inside software you instrument. LangSmith, Langfuse, Arize Phoenix with its OpenInference conventions, Pydantic Logfire and Datadog LLM Observability are among the tools for it. A crew of coding agents is a third subject: the workforce that builds the software, watched while it builds.
| A service at work | An LLM app at work | A coding crew at work | |
|---|---|---|---|
| Unit | A request, milliseconds | A model call or an agent run, seconds to minutes | A work item, minutes to days, across agents, vendors and sessions |
| Failures | Errors, latency, saturation | Answer quality, tool errors, latency, cost | Stalls, drift off the brief, cyber holds, compaction, a model other than the one requested |
| Data | A collector fleet | The code you instrument | The agents' own session files, on the machine that runs them |
| Scarce | Compute, the on-call engineer | Tokens and latency | The operator's attention, the seats' usage windows |
| Who looks | An engineer paged at 3 am | The app's developers | An operator catching up in the morning |
The tools of the second kind are good at what they do, and several now take in Claude Code and Codex, through the agents' own OpenTelemetry export or through hooks that read their session files. Views over weeks are appearing too: Datadog's Agent Console, in preview since June, puts change lead time and pull-request review time next to spend. But in the documentation we reviewed, no tool models a crew: work items with a producer and a reviewer, rounds of review, sprints, messages between agents from different vendors. The work is getting longer as well: METR finds the length of tasks frontier agents can complete doubling about every seven months. And most of these tools were built for code that makes its own model calls. On a seat, the vendor's client makes the call, and what runs out is a usage window rather than a budget.
An open standard, still in the making. The effort to watch is OpenTelemetry's work on generative AI. Its conventions moved into a repository of their own in June 2026 and cover agent spans (creating and invoking an agent, a workflow, a plan), model calls with the requested and the served model side by side, tool calls and MCP. All of it is at Development status, with no release yet, and OpenTelemetry's stability policy advises against long-term dependencies on Development signals. The crew layer is not there yet: tasks and work items (#37), agents messaging agents (#447), long-running work that pauses and resumes (#159), cost (#443) and the providers' safety interventions (#307) are open proposals, and we found none for a subscription's usage windows. The coding agents' own exports (Claude Code, Codex) are off by default and trace a prompt or a turn at a time, not a work item that lives for days.
Why we opened a format of our own, now. We run a multi-vendor crew. On display here are the two key players, Anthropic's Claude Code and OpenAI's Codex, but our meta harness also supports Google Antigravity, and we are adding Kimi and other players soon. We need one record of that crew that every vendor's agents feed, in the crew's own terms: work items, reviews and their rounds, sprints, messages, stalls, drift, holds and seats. That could not wait for the standard, and it should not stay ours either, so we published it as an open format, the Crew Record, under CC BY 4.0. We mean it to sit on top of OpenTelemetry rather than compete with it: its next version is to use OpenTelemetry's names wherever one exists (the gen_ai, messaging and vcs conventions), keep a namespace of its own only for what has none, and read the agents' OpenTelemetry export alongside the harness and the session files. Where our long-horizon data can help the open proposals, we will bring it there, as evidence rather than as a rival convention.
We commit to following that effort. As OpenTelemetry's conventions for agents mature, the Crew Record will move toward them, and what the community adopts from it should end up in the standard itself.
Credits: the ideas in this section belong to the people and projects linked in it, and to the OpenTelemetry contributors writing the open conventions for agents in public. The comparison draws on their documentation as of 27 September 2026.
A lens is only as good as its record. Ours reads one document, the Crew Record: the crew and its roles, the work items and their reviews, the metadata of each message (sender, recipient, kind, size and times; no bodies, prompts or code), and the moments worth finding again. Each row can say whether it is real, derived or synthetic, and a feed that is absent means not recorded, not zero. Today RAIVEN Base Camp writes it from its own logs, and the token and cost figures come from the agents' Codex and Claude Code session logs; the format's next version takes those in directly. It is published under CC BY 4.0, as version 0 with a JSON Schema, so that other harnesses can write it too; adapters for plain Claude Code and Codex are planned next.
The page went live with a mistake in it. It presented the synthetic incident marks as a real crew's record, and it stayed that way for about two hours, until one of four AI reviewers I had set on it (a founder, a client, a growth expert and a platform engineer) flagged it. The sharpest catch of the evening was about the data, not the design. The page says what is synthetic now, down to the 1,068 real messages we moved in time so that they sit inside the work windows they belong to.
Crew Lens runs in our lab today and is being readied for release as an app, not a hosted service: it runs next to your coding agents, reads the record where your harness writes it, and does not send that record to RAIVEN or to anyone else. Your agents are yours, and so is their track record. The coding agents themselves still talk to their vendors, as they do with or without a lens.
The public page walks through it with stills and reels from our record: agenticobservability.app. If you run a crew of coding agents in a harness of your own and want to watch it, we are looking for design partners.
The replay figures were measured with the lens's own time-warp code on the demo's record (27 August to 14 September 2026) at the default speed of sixteen times. The pictures are the lens on that record, except the pet, a screenshot of our desktop; a robot shown stalled marks one of the synthetic stalls. The catch-up figure is the lens's catch-up run on that record, with an absence from 22:30 on 12 September to 08:45 the next morning. The cyber marks come from the agents' session logs: each is a Codex turn that ended with a cyber policy error, bound to the agent that ran it. The costs come from the crew stats, built from the agents' own session logs and priced at the list rates and exchange rate of each turn's day. The 38 to 40 times is the crew's API-equivalent burn over that window, 8,606 to 9,059 euro, against what the two seats cost for the same days, 225 euro.
P.S. New here? What a RAIVEN field report is, and what it refuses to be:
Let's be honest: most AI-generated content in your feed right now is slop. Prose written from other prose, about nothing that happened, worth nobody's time. Here is my promise instead: these posts are AI-polished, but never AI-invented. Every one of them is anchored in real work in the technology trenches.
They are first-hand field reports from RAIVEN, my AI lab in Eindhoven: a working multi-vendor fleet of AI agents that I run daily, both to build real systems and to find out exactly where this technology breaks, and what imperfect but practical remedies push it further. Some reports come with full forensics behind them, down to file hashes. Others describe a pattern from real work whose details stay private: client engagements, security-sensitive systems. Each report says which kind it is, and its claims never pretend to more evidence than it shows.
I write first for executives and decision makers steering in-house AI efforts. The strategy world is full of AI decks assembled by people who never operated the technology they present. These reports are the opposite: trench-level truth, translated into what it means for your decisions. I use my own agents to draft and polish; I am not a native English speaker, and I have a fleet for a reason. But the experiences are real, the judgments are mine, every post is vetted and signed by me, and the numbers are checked against the underlying evidence before publishing.
Field reports are dated observations, not eternal truths, and I will get some things wrong. When a substantive mistake surfaces, whether I find it or a reader does, I publish a dated correction on the post itself. Being corrected in public is part of the method here, not a failure of it.