Skip to main content

Roadmap

Where avuru obs is headed. This is directional, not a commitment — scope and order shift as we learn. The authoritative technical detail lives in the engine repo's ROADMAP.md; this page is the human-readable summary. For what already works today, see Feature status; for what has shipped, see the Changelog.

North star​

:::tip The wedge A fresh Kubernetes cluster → live service map in under five minutes, zero app changes. Every milestone is judged against this, and it's enforced as a CI gate. :::

v0.1 — the wedge (released 2026-07-15)​

The signal tiers we ship for 0.1:

TierSignalStatus
FullService map + RED metrics; trace explorer (waterfall, search)Shipped
BasicLogs (collection, full-text search, trace_id correlation)Shipped
LiteContinuous profiling (per-service CPU flame graphs)Shipped
SupportingInfra metrics (node/pod CPU, memory, network)Shipped

Plus the hard promise: an OTLP drop-in replacement — already-instrumented apps migrate by changing only the exporter endpoint, no SDK or code changes.

Milestones toward v0.1​

M1

Local stack & ingestion

Shipped
  • make dev compose stack (ClickHouse + collector + demo app)
  • OTLP ingest end-to-end; first drop-in e2e test
  • Trace explorer, logs, service map & System Status live
M2

Deployable OTLP backend

Shipped
  • Helm install path; gateway → ClickHouse → hub in-cluster — live
  • Sensor DaemonSet: zero-code eBPF traces + RED (OBI), zero-config logs — live
  • Service inventory UI (RED per service) — live
M3

Signal depth & correlation

Shipped
  • Infra metrics & node/pod health dashboards — live
  • RED metrics dashboard — live
M4

UI depth

Shipped
  • Trace waterfall + flamegraph, spans table, statistics, trace graph — live
  • Trace comparison (structural diff) — live
  • Continuous profiling UI (per-service flame graphs) — live
M5

Gateway build & TTV gate

Shipped
  • OCB-built minimal collector distro — live
  • kind-based time-to-value gate enforcing the sub-5-minute wedge — live

v0.2 — depth and control (released 2026-07-28)​

Everything targeted at v0.2 shipped in v0.2.0: secure-by-default authentication with per-project roles and OIDC SSO, the module framework (pick your signals), error tracking with a browser ingest path, service health groups with criticality tiers, alerting webhooks, network health on the service-map edges, the green energy & carbon module, and a CI-proven "safe to leave on" sensor. The project is licensed AGPL-3.0 as of this release. Full detail on the feature status page and in the changelog.

v0.3 — tenancy you can trust (released 2026-07-31)​

v0.2 secured the read side; v0.3.0 closes the write side. Projects became something you administer — create, rename and delete them from the UI — and per-project ingest keys mean a sender no longer just claims a tenant: in enforce mode the key decides where its telemetry lands. Alongside that: a one-click read-only demo, green energy on RAPL-less cloud VMs, the control-plane groundwork for runtime collection switches, and the rename of the deploy layer from avuruops to avuruobs (breaking — see the upgrade guide). Full detail on the feature status page and in the changelog.

v0.4 — accounts you can administer (released 2026-08-07)​

v0.4.0 finishes the account lifecycle v0.2 began: user management from the UI — edit a user's name and role grants, reset a password, delete an account behind a disable-first rule — plus a self-service password change in Settings → Account, with password operations refused for SSO users whose credential lives at the identity provider. Reviewing that surface closed three ways into an account: a local password that could be minted on an SSO-only account, a login lockout that rotating IP addresses walked straight past, and an SSO login able to take over a local account's email. On the operations side, an install whose schema migration never ran now repairs itself and reports applied-versus-expected schema in Settings → Status; green no longer takes the sensor down on nodes without RAPL; and login works behind a reverse proxy that rewrites Host. Full detail on the feature status page and in the changelog.

v0.5 — operate it from the UI (released 2026-08-17)​

Almost everything you administered needed a values edit and a redeploy; v0.5.0 moves it into the app. Runtime collection control shipped complete: each signal switches on or off from Settings → Collection and the sensor follows in seconds, still default-off behind a narrowly-scoped Role. Service groups are authored in the app, with chart-declared groups read-only and winning name collisions. Settings → Storage and Access show where telemetry lives and who may touch what; the SSO group→role mapping is editable beside the chart's rules; and personal API tokens give scripts and future clients a credential that follows its owner's live permissions. The Dashboard answers "how is the estate doing?" on one screen, the service map carries real health rings, caller-side per-edge latency and shareable filters, and the Nodes screen sorts and filters. OpAMP remains the destination for collection control: status reporting first, remote-config when the upstream grows a client. Full detail on the feature status page and in the changelog.

v0.6 — open at both ends (released 2026-08-22)​

The product observed one cluster well and spoke one protocol; estates are bigger than that. v0.6.0 opened both ends of the pipe. Four more push protocols — Jaeger, Zipkin, Prometheus remote-write and Loki push — land beside OTLP, one values flag each and all off by default, every one through the same tenant stage so ingest keys are enforced whatever the protocol arrived on; forwarding exporters (OTLP/Kafka) dual-write to the backend you already run, behind a bounded queue so a dead target cannot backpressure storage. The claim is a CI gate: real fixtures per protocol through the receivers a real Helm install renders. Member projects let one project read the union of several clusters on every screen, and component toggles (hub.enabled, ui.enabled, gateway.enabled) let a secondary cluster install the ingest half alone against the central store. Projects gained the operational parts they were missing: a retention window of their own and their own storage usage beside the install's. On the energy side, /green now names the nodes its numbers came from and carbon budgets say whether they can actually reach anyone. Full detail on the feature status page and in the changelog.

v0.7 — the clients and the labels (released 2026-08-23)​

v0.6 opened both ends of the pipe; what arrived was only as useful as the words it could be filed under and the surfaces that could read it. v0.7.0 added both. Business tags map a Kubernetes pod label once and carry it onto every signal as avuru.tag.<key>, applied at collection so uninstrumented workloads are tagged too, then offered as a filter on traces and logs — a trace matching when any participating service carries the tag. Declared service metadata lets a service state its own domain, environment and tier as resource attributes and be grouped accordingly across Kubernetes namespaces, with a warning when a declaration cannot be honoured. Two more clients read the same public API: the avuruobs CLI, whose --fail-on predicate gates a deploy with three exit codes, and a Grafana data source — backend-side, so the API token never reaches a browser. Every screen now links to the page of the manual that explains it, and inter-zone traffic accounting reports bytes per availability-zone pair from kernel flows. The map also stopped lying: mesh proxies are recognised as transport and no longer drawn as dependencies. Full detail on the feature status page and in the changelog.

v0.8 — the map grows up (released 2026-08-24)​

The map has been the product's front page since v0.1 and a graph of circles for just as long. v0.8.0 taught it to say more, with no new collection. Virtual targets put the databases, caches and message brokers your services depend on onto the map, derived from the exit spans of the services calling them — a broker drawn from both ends, so a queue is never a dead end. Undetected peers recover the far end of a connection nobody instrumented, which the renderer used to discard outright. Boundaries group the graph by Kubernetes namespace or by service group, resolved exactly as the health board resolves it, and edge volume labels every path at once when the question is which one carries the traffic. The sidebar became layers — Topology, Signals, Operations, Infrastructure — instead of one ever-growing list, with the first-five-minutes path unchanged. Full detail on the service map, the feature status page and in the changelog.

v0.9 — the mesh and the kernel (released 2026-08-25)​

v0.8 stopped the map lying about a meshed cluster. It did not yet tell the whole truth about one: with the proxies hidden, the false dependencies were gone and the real ones behind them were still missing.

The hub now walks each trace's own ancestry across the proxies and reports the dependency underneath, named with the proxy it came through — and the mesh toggle swaps representations instead of stacking them, so a request is never counted twice. The mesh itself gained a screen: per-proxy load, latency and success rate with the calls carried in and out counted apart, plus control-plane health including the configuration your proxies refused.

Underneath, the eBPF sensor moved to a version that exports TCP retransmits — packet loss a fast-looking link can hide — and its per-edge attribution is now asserted on a live kernel in CI rather than assumed. And endpoint checks answer the question observed traffic cannot: whether anything is serving when nobody is calling. Full detail on the feature status page and in the changelog.

v0.10 — what it costs (released 2026-08-26)​

Every release before this one answered what is happening. This one answers a question that arrives from someone who will never open a service map: what is this cluster costing, and how much of that is buying nothing at all.

It answers it from capacity rather than from an invoice. A cluster is sized and billed for what its workloads reserve, so the gap between that and what they draw is the waste — and cost & waste ranks every workload and node by exactly that gap. Idle is measured against the peak a workload reached, never its average, because a request cannot be cut below the peak without inviting the eviction that peak would cause; a workload that declares no request at all is called out as its own state rather than shown as a zero. Rates are values you set, and with none set the screens report cores and bytes and say so — there is no pricing API here, and nothing leaves your cluster to produce the number.

Beside it, the service map stopped depending on what you named things: a gateway called anything at all is recognised from the labels your mesh writes on its own data plane, and those labels only ever promote a workload to transport, never demote one. And a silent control plane now says which of three silences it is — nothing is scraping it, the target is not answering, or it answered with metrics this product cannot read. Full detail on the feature status page and in the changelog.

v0.11 — what was already in your traces (released 2026-08-28)​

v0.10 answered what a cluster costs, and needed new collection to do it. This release adds none at all: every feature in it reads spans the wedge has been storing since the first five minutes, so an install that upgrades sees its history rather than only what arrives next.

The AI module reads the model calls your applications were already sending. Per model: calls, tokens in and out, latency, failures and truncation; per calling service, the same with an owner attached. Four rules keep it honest — the model that answered wins over the one that was asked for, both spellings of the token attributes are accepted, a call that reported no usage is counted and named rather than averaged in as a zero, and a truncated answer is not a failure. Prices are values you declare, absent by default: there is no pricing API here. Alongside it, and deliberately not gated on the module, the gateway drops prompt and completion content by default — a decision about what this product stores, not a feature of a screen you can switch off.

The trace surfaces grew in the same spirit. A Breakdown tab groups spans by service, operation or attribute and weights them by count or by wall time, with a real tail bucket so the parts sum to the whole. A service gained a page of its own, and a single request gained a Path view weighted by the time spent inside each service rather than the time it was on screen. And refused became a third answer beside ok and error — deliberately outside the error rate, because a rejected request is not a broken one. Full detail on the feature status page and in the changelog.

v0.12 — the spend you can act on (released 2026-09-01)​

A report is not an action. v0.12 is about the distance between the two — and it opens by fixing a way of reading spans that turned out to be wrong.

The AI module decided what counted as a model call by testing that gen_ai.operation.name was present, never reading its value. But execute_tool, invoke_agent and create_agent are legal values of that same attribute, so on an agent workload every tool execution was counted as a call to a model: counts inflated, latency quantiles mixed a database lookup with a completion, the model resolved to nothing, and the bucket that exists to name an instrumentation gap filled with spans that were never model calls. Splitting the population by operation class restores all four.

With tools told apart, the rest follows. An agent turn is drawn as the small graph it is — a tool hit four times is one node with a count, because the loop is the thing worth seeing. Spend budgets fire through the alerting you already have, in tokens or money, per calling service or across the estate; a money budget over models you have not priced is refused when the config is parsed, rather than sitting under every threshold by being ignorant of half the spend. And one rate table replaces two formats that could disagree, editable in Settings and applied without a redeploy, with chart-declared values still readable and marked read-only.

The release also carries something that was not planned for it: an MCP server, six read-only tools over the traces, logs, error issues and health an install already stores, so an agent can investigate an incident instead of a person retyping a screen. One handler on the hub — no new collection, no schema, no container — authenticated with the personal API tokens that have existed since v0.5 and resolving their owner's live permissions, and documented in the API reference. It is off by default, because what an agent reads leaves your cluster for whichever model provider you chose: that is a decision to hand over, not one to make for you. Full detail on the feature status page and in the changelog.

v0.13 — dependencies you can see (released 2026-09-04)​

Not a themed line. v0.13 is what was ready, and the two things that were have nothing in common beyond being the follow-through on work an earlier release left standing.

The service page has named callers and dependencies since v0.8, as two tables sorted by volume. A list cannot show shape — that one caller carries all of the traffic, that a single dependency is the one going red, that a hop exists only because the hub walked a trace across a mesh proxy. The Overview tab now opens on a neighbourhood diagram: callers on the left, this service in the middle, what it depends on on the right, each arrow labelled with that path's rate and caller-side p95. It costs no extra request — the page was already reading the service map's response, and the diagram is those same edges drawn. Everything the tables refused to claim, the diagram refuses too: a hop recovered across a proxy reads via <proxy>, an edge nobody timed carries no latency rather than 0ms, and a peer that never sent a span is outlined rather than filled.

The other half is about installing the product at all. v0.12 closed the gateway's Criticals with a collector distro of its own; the node agent could not follow, because it needs contrib-only components and the advisory in question is fixed in no collector release. It gets a distro of its own too, carrying exactly what its rendered config uses and nothing more. Every image the chart pulls by default now scans free of a fixable Critical or High, so a registry policy that blocks Criticals has nothing left to make an exception for — and the exceptions that do remain are named in the values file rather than hidden in a registry. That matters because such a policy surfaces as a rollout that times out, not as a scan report. Full detail on the feature status page and in the changelog.

v0.14 — the estate an agent can reach (released 2026-09-06)​

v0.12 said an agent can read the estate, and it could — over a token a person had to mint and paste by hand, from a server nothing routed, describing a meshed cluster differently from the way the screens describe it. Each of those is a gap between what the release claimed and what an operator got.

A hosted assistant can now sign in. The MCP server speaks OAuth 2.1: discovery, dynamic registration, authorization code with PKCE, rotating refresh tokens. Access tokens are opaque and bound to the MCP endpoint, so one can never be replayed against the rest of the API, and what it may reach is re-read on every request rather than frozen into a signed claim — which is what makes disconnecting take effect on the application's next call. The consent screen states that approving sends traces and log bodies out of the installation, marks the application's self-declared name as unverified, and limits access to one project; Settings → Access lists what you have connected. Off by default, behind its own switch.

And the same estate, described the same way. service_context now reports the dependency behind a mesh proxy rather than the proxy — the collapse the service map has done since v0.9, through the same code rather than a second implementation. The trace Path view reads the hub's per-service self time instead of computing its own, which surfaced a real defect along the way: the MCP tool counted only a raw error status, so a service returning 5xx from an auto-instrumented client was reported as healthy.

A mesh you can actually read. The Mesh screen answered two questions and, on a cluster running without sidecars, rendered two rows. It became a console: every proxy gains a role — control plane, gateway, waypoint, ztunnel, sidecar — and a namespace, read from labels storage was already keeping; two columns that said bytes and showed call counts are renamed, and real bytes and link health take columns of their own; a proxy opens on what it carries, the collapsed dependencies recovered through it, read from the other end. Then a second, separately granted mesh-config module reads the cluster's own Istio and Gateway API objects, read-only, so a namespace that is enrolled and silent finally has a row, and six checks name the misconfigurations that emit no telemetry at all. It is its own module, born off, because the mesh screen needs no cluster permission and a cluster-wide read should be a decision an operator makes rather than one an upgrade makes for them.

The line also carries three repairs. /mcp reached nothing on a Helm install — the endpoint v0.12 announced answered with the UI's 404 page for its whole life. The v0.13 neighbourhood diagram drew a proxy as a caller on a meshed cluster. And release images are cross-compiled rather than emulated, which is what stood between a pushed tag and a published release. Full detail on the feature status page and in the changelog.

v0.15 — what the mesh was told, and what it did (released 2026-09-07)​

v0.14 gave the mesh a console and a reader for its configuration, and left one question standing: the mesh's configuration says a namespace is strict, and a mode is a claim. Whether the traffic was actually encrypted, why a proxy refused a request, and which workloads are really in the mesh are the questions an operator asks first — and every one of them needs a source the product has not read: the proxies themselves, and the pods.

Declared vs observed is the release-defining item. The sensor reads what the data plane reports about itself — per request and per connection, mutual TLS or plaintext — from the proxies on each node, discovered through the annotations the mesh already writes. The hub joins that to the policy that actually governs each workload, including the selector-scoped ones the v0.14 reader skipped. Every namespace, workload and map edge gets a lock; "permissive with plaintext callers" names the callers; and "declared strict, observed plaintext" — a policy that is not applied — becomes a finding instead of a silence. Under the mesh module, on by default: the module is the consent.

Around it: every workload, in or out — the configuration reader adds pods, read-only and capped, so a Workloads tab lists what the cluster runs whether or not it sent traffic, with whether a sidecar is in the pod, whether the node agent captured it, which waypoint binds it and which policies decide its mTLS mode. Six checks become seventeen, including the two v0.14 held back for lack of pods and the ones a security posture needs, and every pod-dependent check says so when the pod list was cut, so an empty issues column is never a clean bill. The proxy explains its failures — requests by response flag and by destination version, with counters a default mesh does not expose named as not collected rather than rendered as zero. And the graph reads by role, with a marker at the caller end of every measured edge for all-encrypted, mixed and plaintext, on the map and the mesh graph alike.

Shipped as the v0.15.0 release; the release-defining item has its own post. Deliberately not in this line: per-proxy effective configuration, which stays a debugger's job; posture history and alerting on posture, which wait until the verdicts have been trusted on real clusters; and a second control plane, which still waits on an operator running one.

v0.16 — the workload's page (released 2026-09-07)​

v0.15 put, on one row, what the mesh was told and what it did. On a real ambient cluster with a mesh console beside it, the reader's next questions were the ones that row did not answer: what is this workload, since when, which pods on which rollout, what configuration names it — and what did the proxies write about it. The record was already in objects the reader watched; the proxies' lines were already in the store, filed under the proxies' own names. Nothing new was read from the cluster.

The record is the release-defining item. A workload's page carries the controller's creation time (or the oldest pod's, and it says which), its type, app and version, every label and the controller's annotations within stated bounds, each pod with the rollout it belongs to, and a health verdict with the reason that decided it. Beside the policies that select it by label, the routes and rules that reach it through its Services — HTTPRoute, GRPCRoute, VirtualService, DestinationRule — resolved through the same host index the checks use, each with its own findings. A routed workload is no longer called unconfigured.

Three sources, one stream. The application logs one side of every request; ztunnel logs the connection and the waypoint logs the HTTP exchange, under their own service names. One route reads all three for a workload as one ordered stream under one keyset cursor — composed in the hub, which knows the pods behind the workload and the waypoint it is bound to. When the pods cannot be known, the proxies' lines are matched by name and namespace together, and the response says which rung it fell to.

Shipped as the v0.16.0 release. Deliberately not in this line: pod readiness, restarts and container states — the reader projects a dozen fields per pod on purpose; and a ztunnel row on the Proxies tab seeded from the scrape rather than from spans, which waits for the next line with the TCP byte counters the scrape already keeps.

v0.17 — the connected workspace (released 2026-09-15)​

The map is the application entry point. Explorer keeps a selected service's callers and dependencies beside the graph, then links into its traces, logs and service detail. Selection travels in the URL and is available from the keyboard. An empty project teaches eBPF and OTLP connection; a failed read offers retry. The shared UI adopts sage and forest surfaces with mobile navigation that keeps enabled modules and project controls. The Errors screen gains a summary over the matching issue set, a service's Logs tab resolves its workload, log tables copy and download, and container log bodies yield a severity.

Shipped as the v0.17.0 release.

v0.18 — look inside the cluster (released 2026-09-16)​

Infrastructure gains Cluster X-Ray: an interactive isometric view of Nodes, translucent Pods and a logical infrastructure layer beside the inventory. Orbit, zoom, separate or hide layers, filter namespaces, inspect a Pod's CPU, memory and placement, and follow the connections the traces recorded — only between recorded Pod identities, with unresolved addresses kept explicitly unlocated. The renderer loads on demand, the scene states its limits, and the inventory stays for keyboards and browsers without WebGL.

Shipped as the v0.18.0 release.

v0.19 — several services' logs as one stream (released 2026-09-18)​

One log explorer for Signals → Logs and Service Mesh → Logs: pick services and workloads, choose which sources to read, and follow them merged newest-first or as one panel per service. Selection, sources and display live in the URL; GET /api/v1/logs composes the subjects in one paginated query under a signed pagination token. The dark theme moves to slate with readable secondary text. v0.19.1 narrowed the mesh screens to the namespaces the caller's projects reach.

Shipped as the v0.19.0 release.

v0.20 — typed decisions, on paper (released 2026-09-23)​

No product change. A design note ranks where a classifier that answers with a calibrated probability rather than text would fit — the severity a rule could not read, errors a fingerprint kept apart, alerts to triage — and what each would cost the promise that nothing leaves the cluster: at most an opt-in, batch module sending structured or already-scrubbed payloads, born off, never the ingest path. An offline harness measures the first case, log-severity backfill, from a laptop against a copy of the log table. The note stays a draft until that measurement runs.

Shipped as the v0.20.0 release.

Beyond v0.20 (directional)​

  • The proxies from the scrape, not from spans. On an ambient cluster ztunnel emits no span, so the Proxies tab has no row for it even while its metrics are read; the scrape's up and gauges should seed the row, and the TCP byte counters already scraped should be read.

  • Read a second control plane, once someone running one can say which of its signals answer the questions the Istio card answers — and which simply have no answer there.

  • Cost joined to green: the same reserved-and-idle capacity in Wh and gCO2e, on an install running both. One story about waste, in two units.

  • The incident: root-cause summaries, with any provider switch disabled by default — it would be the first outbound call in a product whose whole promise is that nothing leaves the cluster. v0.20's design note takes the decision half — how urgent, and is it a symptom of an upstream alert already firing — onto paper as a typed decision over structured metadata, with nothing generated; the summary half stays out.

  • Scripted multi-step check journeys, if demand appears — v0.9 ships single-request checks deliberately.

  • Deeper profiling: off-CPU and memory profiles, as the upstream eBPF profiler grows them.

  • Storage re-evaluation: ClickHouse stays behind the storage.Store interface; GreptimeDB is slated for re-evaluation mid-2027 without changing hub code.

How this roadmap changes​

Open an issue or discussion to propose a change of direction; larger items graduate into an Avuru Enhancement Proposal before implementation. Roadmap edits go through a normal pull request.