RAIVEN Field Report · Agentic observability

Making the night shift watchable

A crew of long-horizon coding agents works through the night and leaves a record that is tedious to read. Our side quest: agentic observability that people will actually watch. Dare I say, less boring.

Richard Vdovjak · RAIVEN 2026-09-27 AI-polished, never AI-invented · anchored in real work in the technology trenches. Read more about our promise.
The record of one agent in Crew Lens, x3: its robot and the models it was served, beside cards for its days of service, token burn, projected cost at API prices, token mix, producing and reviewing time, where its model time went, its last health check, compactions, messages, and its stalls, drift and cyber holds.
One agent's record in Crew Lens: x3, a Codex agent, over the 18 days. 2.91 billion tokens, about 1,665 euro at API prices.

TLDR, for decision makers

Agent avatars are becoming a thing, and they are making their way into developer culture. The vendors have their own take on it. OpenAI's Codex app has pets: small animated companions that follow the work, picked from a built-in set or hatched to your own design. Anthropic gave Claude Code a small creature that watched you code, for a while (an April Fools' joke, since retired). Here is ours: the rAIven, our Codex pet, perched on the bespoke gauges we built to watch our two seats.

A metallic raven with a glowing blue eye, our Codex pet, perched on two gauges of our own: the ChatGPT seat at 36 percent of its weekly allowance, and the Claude seat at 42 percent overall, 0 percent for Fable and 5 percent of its five-hour window.
The rAIven, our Codex pet, on our seat usage gauges.

Frankly though, these avatars do little beyond being cute and offering a few UX shortcuts. A pet shows whether an agent is running, waiting for you or blocked, right now, in one vendor's app. That does not help with the observability problems of long-horizon work by a crew of coding agents, and even less when the crew is mixed from several vendors. The idea of an avatar may still be a useful one: it can make the otherwise mostly boring work of reading logs and traces more fun. So we improved on the avatar, adding graphic metaphors for things like the agent's ID, its model and reasoning effort, and its context window (size and fill), which make an agent avatar more useful at a glance.

Our agent avatar building from its wireframe under a green scanning plane on a pedestal, above its controls: vendor, model, effort, context, age, rebuild and turntable.
Our agent avatar, live: turn it, and change its vendor, model, effort, context and age. It loads from agenticobservability.app, where the agent viewer has more.

Our lab runs a crew of AI coding agents. Not a chatbot and not a copilot: a crew. In the stretch of record behind this report, late August to mid September 2026, it was five Claude agents and twelve Codex agents, plus one human operator: me. They work in RAIVEN Base Camp, our multi-vendor meta harness. They take work items, build, review each other's code across vendors (the xLLM review from an earlier report), run sprints under a sprint master, and message each other and me. Much of the work is long-horizon: an agent carries a task across hours, hundreds of steps and more than one context window, often overnight. The projects are RAIVEN's own, so nothing shown here holds client data.

What a crew leaves behind: the raw records, drifting past on slabs, as on agenticobservability.app.

A record few people read

A crew like that leaves a record. Work items and their reviews, the messages between agents, model changes, context compactions, sub-agent fan-outs, security warnings. It is the audit trail (which agent changed what, on which model, at whose request) and the first place to look when something breaks. Over weeks it tells you how the workflow runs in practice, and which models and settings suit which jobs.

It is also thousands of events a month, with far more log lines behind them. Reading through them is hardly a dream job. Dashboards help, and we have them, but a dashboard is a snapshot: it shows where things stand, not how they got there. And a log file is not something you put in front of a board, a client, or a friend who asks what exactly the agents did last night.

Agentic observability, for a crew

The phrase is getting crowded. Tracing a single model call or a single agent run is useful, and tools for it exist. Our questions were about the crew: who is stuck, what it cost, what is about to ship; which review is missing because a vendor's classifier stopped it; whether the model an agent was served is the one it asked for; what happened while I was asleep. Those questions cut across vendors, sessions and weeks, so the answers need the crew's whole record in one place, in time order. Observability itself is an established discipline; what we borrow from it, and where a crew differs, is further down.

The side quest: a theme

So we made that record watchable. Crew Lens is a web-based 3D visualisation library that runs next to the coding agents. Its first theme, the outpost, draws the crew as a camp in the mountains. Each agent is a small robot, and the things on its body are data: the bubble round its head is the model and its reasoning effort, the tank on its back is the context window (it fills as the agent works, and at the top the agent compacts), the mark on its neck is the vendor, the chevrons on its arms are days of service. A work item is a black rock the agent chips open until a diamond shows; converged work is a green diamond on the sprint's pile, scrapped work a black one, kept for the record. Each sprint gets a hangar, sized by the code it touches. A second theme draws the same record as craft building a station in orbit.

The outpost at 17:00 on 4 September, replayed from the record: glass hangars under sunset mountains, green diamonds of finished work on the sprints' piles, a few agents at their benches, the rest of the crew as robots standing on their pads, and the operator hovering above the camp.
The outpost at 17:00 on 4 September, replayed from the record: six agents at work, the rest on their pads.

Why robots? Because attention is the scarce resource. Long-horizon agentic work is interesting to the people who run it and hard to explain to anyone else. I wanted a way to talk about it with people who will not open a log file, and a way to look at my own crew that I would pick over the alternative.

A customisable event timeline is where Crew Lens really shines

Some way into the side quest, I started to suspect that the robots were not the most useful part. The timeline under them was. The lens plays the record back like a film, and above the scrubber runs an event timeline: one row per kind of moment worth finding again. Sub-agent fan-outs, stalls, drift flagged by our Sentinel watchdog, cyber classifier warnings and holds, a served model that differs from the one requested, review convergences, and my own nudges as operator. Click a mark and the replay jumps to five seconds before it, just in time to watch it happen.

The lens's timeline flags, jump to: fan-outs off, stalls on, Sentinel drift on, cyber classifier (warnings, holds) on, served model not equal to requested off, convergences off, operator nudges on, and incident replay only (skip the mundane) on.
The timeline's flags: pick the kinds of moment to mark, and replay only those.

Two things shrink weeks to hours. The replay squeezes the quiet: much of an agent's day is waiting, for a review, an answer or the next job, so once three minutes pass with nothing happening, the rest of the gap plays in fifteen seconds of replay time, whether it lasted ten minutes or a night. And incident replay plays only the kinds of moment you ticked, from eight minutes before each to twenty minutes after, and jumps over the rest.

Measured on 18 days of our crew's record at the default speed of sixteen times: the whole record plays in 26.8 hours; skipping the quiet brings it to 7.5 hours; the incidents alone (stalls, drift, cyber, model mismatches) take 1.8 hours, and the cyber classifier's events about 7 minutes.

Eighteen days, under two hours of trouble. Mostly synthetic trouble, for now.

How long 18 days of the crew's record take to watch at the replay's default speed: 26.8 hours as it happened, 7.5 hours skipping the quiet, 1.8 hours for the incidents only, 7 minutes for the cyber classifier's events only. Below, the timeline filtered to those incidents.
18 days, four ways to watch them. Below, the timeline filtered to the incidents (real cyber marks; the other incident marks synthetic).

A caveat on those numbers: the stalls, drift and model mismatches in that record are synthetic data, and so are its fan-outs and most of the operator's nudges, generated because the feeds for them are not wired into the record yet. Each of those kinds of event is real and observable in our crew. The cyber marks are real: we read them from the agents' session logs. The figures show how the replay behaves, not how often our crew stalls or drifts.

Show me what I missed

The crew keeps working while I sleep, sit in a meeting or take a weekend. The lens notices when its view goes unattended: fifteen minutes without a pointer move, key, scroll or touch counts as away. Come back, and the timeline shades the stretch you missed, a bar sums up what happened in it, and Catch up plays only the flagged moments inside it, at a speed that keeps the whole absence to about 45 seconds, then returns to live. On the night of 12 September the bar reads three cyber holds, all on one agent, the longest from 02:23 to 07:59; the catch-up plays them in about four seconds. Watching counts and jumping does not: play a stretch through and it is marked seen; click about in it and it stays for later.

Back in the morning: a bar over the timeline reads Away 10 h 15 min, with three cyber holds, and offers a catch-up of about 4 seconds; below, the catch-up playing through the night camp, the held agent lying at its bench under the screen of its hold.
Back in the morning: ten hours away, three real cyber holds on one agent, a four-second catch-up.

Price tags

Then I wanted the money on the scene. With price tags on, each finished diamond carries what its work cost, in euro at API list prices, or in tokens, and each hangar sums its pile so far, scrapped work included, so the tags grow as a replay runs. The cost of an item is attributed by time (each model turn is split evenly across the items its agent held open at the time): an estimate, not an invoice, and how close it lands has not been measured yet. On our largest sprint's hangar, the 36 finished items with attributed turns came to about 971 euro at API prices, the dearest of them to about 155 and the scrapped ones to about 85.

The outpost with price tags on: the largest sprint's hangar reads 46 finished, 36 of them priced, about 971 euro, scrapped about 85 euro, and each priced diamond on its pile carries its own tag.
The largest sprint's hangar in the demo's record: 46 finished items, the 36 with attributed model turns about 971 euro at API prices, 85 of it on scrapped work. Each priced diamond carries its own tag.

The crew does not pay API prices. It runs on two flat-rate seats, a Claude Max 20x and a ChatGPT Pro 20x plan, inside the vendors' terms: each agent is an ordinary session of the vendor's own client, Claude Code or Codex, on our own seats, with no borrowed credentials and nothing resold. Priced at API list rates, its burn over those 18 days comes to 38 to 40 times what the two seats cost. That multiple is not a return on investment (tokens are not outcomes); it is the seat economics of our earlier report, now visible per item, per sprint and per agent.

The crew board's seat cards for the 18 days: the two seats cost 225 euro, against 38.2 to 40.2 times that in API value; the Claude Max 20x seat, 72 percent of its burn from the crew, 3,230 to 3,683 euro at API prices against a 110 euro subscription; the ChatGPT Pro 20x seat, 44 percent, 5,376 euro against 116.
The seat cards on the crew board, for the same 18 days.

The false positive that costs a review

One row on the timeline earned its own section. When a model vendor's cyber classifier flags an agent's work, the two vendors in this record react differently. Codex ends the flagged turn with a policy error and the agent stops, which the scene draws as a robot under a blue screen. The block may lift within minutes or last hours, and in our logs the stops come in episodes over days. Claude Code does not stop. It shows a notice, re-runs the request on an older model and stays on that model until someone switches it back by hand. For an unattended crew, that is an agent working below the model I chose until I happen to notice.

Something to consider

Claude Code has a setting for this: switchModelsOnFlag, or "Switch models when a message is flagged" in /config. Turn it off, and a flagged request pauses the session with two options: switch to the fallback model, or edit the prompt and retry on the current one (Claude Code docs).

Without observability, I would take the pause. A full block is hard to miss: it costs time at once, but I know about it. A drop in model quality can go unnoticed and bite much later, when the project is far ahead; late bugs take longer to fix, and the end users may meet them too.

With observability in place, there is a better option than a stop: let the agent carry on, mark the flagged stretch, and add an extra review for the work done in it.

The cyber marks in the public record are real, read from the agents' session logs: six holds and two warnings in those 18 days, all on Codex agents. The longest held one agent overnight, from 02:23 to 07:59, which is the kind of night the catch-up is for.

Before dawn on 13 September: x6 lies beside its pad next to a blue screen that reads cyber hold, while the rest of the crew stand on their pads.
x6 at 07:37 on 13 September, more than five hours into a real cyber hold.

This has had my attention since our earlier report, the open proposal on why the gate fires on a crew's own work and how the vendors could make it more precise. Since then I have also tried to keep ordinary reviews out of the classifier's way, and made that explicit: I ask the agents not to do adversarial review work on ordinary coding, such as UX. Before that, they often volunteered an adversarial take on any review task, however ordinary.

That habit became counterproductive. Each hold stopped an agent reviewing our own code, already written, and on our crew each hold so far has been reviewed and judged a false positive. A review that stops can miss the bug it was there to find, and a bug nobody found stays unpatched. So a false positive can leave our code less secure, not more, the opposite of what the vendors intend. It also leaves a missing review that the crew may be waiting on. The lab's full record of these holds, back to June, deserves a report of its own.

So the next things the lens learns: to say which review is missing ("review incomplete: held") and suggest handing it to an agent from another vendor; and to mark the stretch a Claude agent spends on the fallback model, and send the work done in it for an extra review.

Observability is an established disciplineWhat we borrow from it, where a crew of coding agents differs, and why we opened a format of our own

None of this starts from scratch. The word comes from control theory: Rudolf Kálmán defined observability in 1960 as how well a system's internal state can be inferred from its outputs (On the general theory of control systems). Charity Majors and her colleagues at Honeycomb brought it to software with a practical test: can you understand whatever state your system has got into, from the outside, without shipping new code (Honeycomb)? Monitoring watches for the failures you expected; observability lets you ask the questions you did not think of in advance.

The practice around it is mature. Peter Bourgon's metrics, tracing and logging became the three pillars, and Ben Sigelman's critique of them, with Brandur Leach's canonical log lines at Stripe, made the case for wide structured events: one dense record per unit of work, carrying every dimension you might filter on. Google's site reliability engineering books set out service level objectives, error budgets and percentiles over averages, and objectives for data pipelines that judge freshness and completion rather than latency. John Allspaw made the case for the blameless postmortem at Etsy. DORA's research programme measures software delivery, and its 2025 report calls AI an amplifier of a team's existing strengths and weaknesses. And OpenTelemetry, formed in 2019 from OpenTracing and OpenCensus, graduated in the Cloud Native Computing Foundation in May 2026 as the vendor-neutral way to produce, carry and name telemetry; its wire protocol, OTLP, is stable for traces, metrics and logs.

Three things to observe. Classic observability watches a service while it serves. The newer branch, LLM observability, watches an application that calls models: it traces each model call, tool call and agent run inside software you instrument. LangSmith, Langfuse, Arize Phoenix with its OpenInference conventions, Pydantic Logfire and Datadog LLM Observability are among the tools for it. A crew of coding agents is a third subject: the workforce that builds the software, watched while it builds.

A service at workAn LLM app at workA coding crew at work
UnitA request, millisecondsA model call or an agent run, seconds to minutesA work item, minutes to days, across agents, vendors and sessions
FailuresErrors, latency, saturationAnswer quality, tool errors, latency, costStalls, drift off the brief, cyber holds, compaction, a model other than the one requested
DataA collector fleetThe code you instrumentThe agents' own session files, on the machine that runs them
ScarceCompute, the on-call engineerTokens and latencyThe operator's attention, the seats' usage windows
Who looksAn engineer paged at 3 amThe app's developersAn operator catching up in the morning

The tools of the second kind are good at what they do, and several now take in Claude Code and Codex, through the agents' own OpenTelemetry export or through hooks that read their session files. Views over weeks are appearing too: Datadog's Agent Console, in preview since June, puts change lead time and pull-request review time next to spend. But in the documentation we reviewed, no tool models a crew: work items with a producer and a reviewer, rounds of review, sprints, messages between agents from different vendors. The work is getting longer as well: METR finds the length of tasks frontier agents can complete doubling about every seven months. And most of these tools were built for code that makes its own model calls. On a seat, the vendor's client makes the call, and what runs out is a usage window rather than a budget.

An open standard, still in the making. The effort to watch is OpenTelemetry's work on generative AI. Its conventions moved into a repository of their own in June 2026 and cover agent spans (creating and invoking an agent, a workflow, a plan), model calls with the requested and the served model side by side, tool calls and MCP. All of it is at Development status, with no release yet, and OpenTelemetry's stability policy advises against long-term dependencies on Development signals. The crew layer is not there yet: tasks and work items (#37), agents messaging agents (#447), long-running work that pauses and resumes (#159), cost (#443) and the providers' safety interventions (#307) are open proposals, and we found none for a subscription's usage windows. The coding agents' own exports (Claude Code, Codex) are off by default and trace a prompt or a turn at a time, not a work item that lives for days.

Why we opened a format of our own, now. We run a multi-vendor crew. On display here are the two key players, Anthropic's Claude Code and OpenAI's Codex, but our meta harness also supports Google Antigravity, and we are adding Kimi and other players soon. We need one record of that crew that every vendor's agents feed, in the crew's own terms: work items, reviews and their rounds, sprints, messages, stalls, drift, holds and seats. That could not wait for the standard, and it should not stay ours either, so we published it as an open format, the Crew Record, under CC BY 4.0. We mean it to sit on top of OpenTelemetry rather than compete with it: its next version is to use OpenTelemetry's names wherever one exists (the gen_ai, messaging and vcs conventions), keep a namespace of its own only for what has none, and read the agents' OpenTelemetry export alongside the harness and the session files. Where our long-horizon data can help the open proposals, we will bring it there, as evidence rather than as a rival convention.

We commit to following that effort. As OpenTelemetry's conventions for agents mature, the Crew Record will move toward them, and what the community adopts from it should end up in the standard itself.

Credits: the ideas in this section belong to the people and projects linked in it, and to the OpenTelemetry contributors writing the open conventions for agents in public. The comparison draws on their documentation as of 27 September 2026.

Labelled data, open format

A lens is only as good as its record. Ours reads one document, the Crew Record: the crew and its roles, the work items and their reviews, the metadata of each message (sender, recipient, kind, size and times; no bodies, prompts or code), and the moments worth finding again. Each row can say whether it is real, derived or synthetic, and a feed that is absent means not recorded, not zero. Today RAIVEN Base Camp writes it from its own logs, and the token and cost figures come from the agents' Codex and Claude Code session logs; the format's next version takes those in directly. It is published under CC BY 4.0, as version 0 with a JSON Schema, so that other harnesses can write it too; adapters for plain Claude Code and Codex are planned next.

The page went live with a mistake in it. It presented the synthetic incident marks as a real crew's record, and it stayed that way for about two hours, until one of four AI reviewers I had set on it (a founder, a client, a growth expert and a platform engineer) flagged it. The sharpest catch of the evening was about the data, not the design. The page says what is synthetic now, down to the 1,068 real messages we moved in time so that they sit inside the work windows they belong to.

Where it runs

Crew Lens runs in our lab today and is being readied for release as an app, not a hosted service: it runs next to your coding agents, reads the record where your harness writes it, and does not send that record to RAIVEN or to anyone else. Your agents are yours, and so is their track record. The coding agents themselves still talk to their vendors, as they do with or without a lens.

The public page walks through it with stills and reels from our record: agenticobservability.app. If you run a crew of coding agents in a harness of your own and want to watch it, we are looking for design partners.

Where the numbers come from

The replay figures were measured with the lens's own time-warp code on the demo's record (27 August to 14 September 2026) at the default speed of sixteen times. The pictures are the lens on that record, except the pet, a screenshot of our desktop; a robot shown stalled marks one of the synthetic stalls. The catch-up figure is the lens's catch-up run on that record, with an absence from 22:30 on 12 September to 08:45 the next morning. The cyber marks come from the agents' session logs: each is a Codex turn that ended with a cyber policy error, bound to the agent that ran it. The costs come from the crew stats, built from the agents' own session logs and priced at the list rates and exchange rate of each turn's day. The 38 to 40 times is the crew's API-equivalent burn over that window, 8,606 to 9,059 euro, against what the two seats cost for the same days, 225 euro.

What this doesn't claim

P.S. New here? What a RAIVEN field report is, and what it refuses to be:

About these field reports

Let's be honest: most AI-generated content in your feed right now is slop. Prose written from other prose, about nothing that happened, worth nobody's time. Here is my promise instead: these posts are AI-polished, but never AI-invented. Every one of them is anchored in real work in the technology trenches.

They are first-hand field reports from RAIVEN, my AI lab in Eindhoven: a working multi-vendor fleet of AI agents that I run daily, both to build real systems and to find out exactly where this technology breaks, and what imperfect but practical remedies push it further. Some reports come with full forensics behind them, down to file hashes. Others describe a pattern from real work whose details stay private: client engagements, security-sensitive systems. Each report says which kind it is, and its claims never pretend to more evidence than it shows.

I write first for executives and decision makers steering in-house AI efforts. The strategy world is full of AI decks assembled by people who never operated the technology they present. These reports are the opposite: trench-level truth, translated into what it means for your decisions. I use my own agents to draft and polish; I am not a native English speaker, and I have a fleet for a reason. But the experiences are real, the judgments are mine, every post is vetted and signed by me, and the numbers are checked against the underlying evidence before publishing.

Field reports are dated observations, not eternal truths, and I will get some things wrong. When a substantive mistake surfaces, whether I find it or a reader does, I publish a dated correction on the post itself. Being corrected in public is part of the method here, not a failure of it.