Releases
avuru obs follows semantic versioning (vX.Y.Z). This page
is the version-level timeline; for granular, dated changes see the
Changelog, and for what's next see the Roadmap.
:::note Pre-1.0 Until v1.0.0, minor bumps may include breaking changes; patch bumps are fixes only. :::
The main trunk. See the Roadmap.
Typed decisions, on paper first. A design note ranks where a classifier that answers with a calibrated probability rather than text would fit — log severity a rule could not read, errors a fingerprint kept apart, alerts to triage — and what each would cost the promise that nothing leaves the cluster. An offline harness measures the first case from a laptop, never from the hub. No product change ships.
See the changelog and the release artifacts.
Fix: the mesh screens respect the project. A project-scoped account now sees only the namespaces its projects reach on Service Mesh; an identity that may see every project still sees the cluster whole.
See the release artifacts.
Several services' logs as one stream. The log explorer
selects services and workloads, reads application, ztunnel, waypoint and other
sources independently, and shows them merged or as panels per service — from
Signals and from a new Service Mesh tab. GET /api/v1/logs composes subjects
with a signed pagination token; GET /api/v1/logs/services suggests them. The
dark theme moves to slate with readable secondary text.
See the changelog and the release artifacts.
Look inside your cluster. Cluster X-Ray adds interactive isometric Node/Pod placement, transparent layers and Pod inspection. Connections use recorded span identities; unresolved addresses remain unlocated. The renderer loads on demand and the inventory remains available without WebGL.
See the changelog and the release artifacts.
Explorer. The service map becomes the application entry point, with an adjacent inspector and links to related signals. Error statistics and service-aware log sources improve investigation.
The workload's page. v0.15 put, on one row, what the mesh was told and what
it did; this release answers the questions that row did not. A workload on the
mesh screen opens on the cluster's own record of it —
created when and by which controller, type, app and version, every label,
the controller's annotations, each pod with the rollout it belongs to, a health
verdict with the reason that decided it — and lists, beside the policies that
select it, the routes and rules that reach it through its Services, each with
its own findings. Then a Logs tab reads the workload's own lines together
with the ztunnel lines naming its pods and the waypoint lines naming its
Service, as one stream under one cursor, composed by the hub that knows the
pods — and says so when it could only match them by name.
No new permission and no new collection: the record comes from objects
mesh-config already watched, the logs from tables the logs module already
fills. Installs without the mesh module are untouched.
See the changelog entry and the release notes.
What the mesh was told, and what it did. The namespace list said STRICT,
and a mode is a claim: a strict policy that is not applied to a workload looks,
from configuration alone, exactly like one that is. This release reads the two
sources the product had not read — the proxies themselves, and the pods — and
puts them on one row. Each node's sensor scrapes the sidecars, waypoints,
gateways and ztunnel on that node, discovered from annotations the mesh already
writes; the mesh-config module reads pods, still read-only, so every
workload the cluster runs has a row whether or not it ever sent a span. A
Security tab draws a posture per
workload — strict and all mutual TLS; declared strict, observed plaintext, which
is a finding; safe to tighten; plaintext callers, named — and the
service map marks every edge a proxy measured. A
Workloads tab lists enrolment as the pods report it, six configuration
checks become seventeen, a proxy's page explains its failures by response flag
and destination version, and the mesh graph draws each role as its own shape.
The data-plane scrape is on by default under the mesh module and starts at
upgrade; mesh.dataPlane.enabled=false keeps the screen without it. Installs
without the mesh module are untouched.
See the changelog entry and the release notes.
The estate an agent can reach. v0.12 said an agent could read the estate,
and it could — over a token a person had to mint and paste by hand, from a
server nothing routed, describing a meshed cluster differently from the way the
screens describe it. This release closes each of those gaps. The MCP server
speaks OAuth 2.1, so a hosted assistant signs in the way
every other client does, on a consent screen that says what leaves your
installation and can be taken back from Settings → Access; tokens are opaque and
bound to /mcp, so one is refused on the rest of the API. service_context
reports the dependency behind a mesh proxy rather than the proxy, through the
same code the service map uses, and per-service
self time is computed once for both an agent and the trace Path view — which
surfaced that the tool had been reporting a service returning 5xx as healthy.
The Mesh screen became a console: every proxy gains a
role and a namespace, real bytes and link health beside its calls, and a page of
its own showing what it carries. A separately granted, read-only mesh-config
module reads your cluster's Istio and Gateway API objects, lists the namespaces
that are enrolled and silent, and runs six checks for breakage that emits no
telemetry.
Three repairs: /mcp is now reachable on a Helm install, a meshed service no
longer shows its proxy as a caller in the neighbourhood diagram, and release
images are cross-compiled rather than emulated — the hour-long arm64 build that
stood between a tag and a published release is gone.
See the changelog entry and the release notes.
Dependencies you can see. A service page could name what calls a service and
what it calls; it could not show the shape of it. The Overview tab now opens on a
neighbourhood diagram — callers on the left,
the service in the middle, its dependencies on the right, each arrow labelled
with that path's rate and caller-side p95. It costs no extra request, because the
page was already reading those edges to fill two tables, and it claims nothing
the tables refused to claim: a hop recovered across a mesh proxy reads
via <proxy>, an edge nobody timed carries no latency rather than 0ms, and a
peer that never sent a span is outlined rather than filled. The map link that
used to strand a service on an empty graph now lands on its one-hop
neighbourhood.
The release also finishes v0.12's work on images. The node agent gets a
first-party collector distro of its own — it needs receivers the minimal
distro does not carry, and CVE-2026-56854 is fixed in no collector release, so
no version bump could have closed it. Every image the chart pulls by default now
scans free of a fixable Critical or High, which means a registry with a
block-on-critical policy has nothing left to except: relevant because such a
registry stops serving a flagged image, and you meet that as a rollout timing out
rather than as a scan report.
See the changelog entry and the release notes.
The spend you can act on — and an estate an agent can read. AI spend stops being a number on a screen: monthly budgets in tokens or money fire through the alerting channels you already have, and model prices and compute rates now live in one rate table with one currency, editable in Settings and applied without a redeploy. The budget evaluator and the screens read the same resolver, so a budget can never be measured against a price other than the one on display.
The AI module also stops counting tool executions as model calls — a defect that inflated call counts, mixed a database lookup into completion latency and filled the no-usage bucket with spans that never called a model. Tools get a table of their own, and an agent turn is drawn as the graph it is.
New in this release, the MCP server: six read-only tools over the traces, logs, error issues and health you already store, so an agent can investigate an incident directly. It adds no collection and no container — one handler on the hub, authenticated with personal API tokens — and it is off until you turn it on, because what an agent reads leaves your cluster.
Both collector images an install runs were also brought below their advisories: the gateway distro pins its dependency floors, and the node agent moved to the collector line that scans clearest.
See the changelog entry and the release notes.
What was already in your traces. No new collection at all: every feature reads spans avuru obs has been storing since your first five minutes, so an upgrade shows you your history rather than only what arrives next. AI observability reports the model calls your applications were already sending — per model and per calling service, with tokens, latency, failures and truncation, and a cost if you declare your rates. The same release takes a position on message content: prompts and completions are now dropped at the gateway by default, because they were being stored under your ordinary retention and shown to every Viewer, and nobody had ever decided that.
Traces gain a Breakdown — the traffic as a treemap and a donut, grouped by anything a span carries, weighted by count or by total time, with a tail that is a real bucket so the parts sum to the whole. There is a page per service, a Path view drawing the service-level shape of a single request, and Refused: server-side 4xx as its own outcome, deliberately kept out of the error rate.
See the changelog entry and the release notes.
What it costs. A cluster is sized by what its workloads reserve, not by what they use — and only the second half of that was ever visible. Cost & waste adds the first: every workload's reserved CPU and memory against what it actually drew, ranked by the gap, with workloads that declare no request at all called out as their own state and node allocation shown beside node usage. Idle is measured against the peak, never the mean, and rates are yours to declare — there is no pricing API and nothing leaves your cluster to produce a number. The service map now recognises a gateway you named anything at all, from the labels your mesh writes on its own data plane, with names still the answer for sidecars. And a silent control plane says which of the three silences it is, one of which is "this is not a control plane we can read".
The mesh and the kernel. On a meshed cluster every call is intercepted, so a dependency graph either draws hops that are not dependencies or draws nothing at all. The hub now walks each trace's own ancestry across the proxies and reports the dependency underneath, named with the proxy it came through — and the mesh toggle swaps representations rather than stacking them. The mesh itself gained a screen: per-proxy load, latency and success rate with calls carried in and out counted apart, plus control-plane health including the configuration your proxies refused. The eBPF sensor moved to a version that exports TCP retransmits, and its per-edge attribution is now asserted on a live kernel rather than assumed — which uncovered a latent crash that took the whole sensor down whenever TCP stats were enabled. Finally, endpoint checks answer the question observed traffic cannot: whether anything is serving when nobody is calling.
The map grows up. Virtual targets put the databases, caches and message brokers your services depend on onto the map, derived from the exit spans of the services calling them — no agent in the dependency, and a broker drawn from both ends. Undetected peers recover the far end of a connection nobody instrumented, which the renderer used to discard. Boundaries group the graph by Kubernetes namespace or by service group, and edge volume labels every path at once. The sidebar became layers — Topology, Signals, Operations, Infrastructure — with the first-five-minutes path unchanged. All of it from telemetry already arriving: no new collection.
The clients and the labels. Telemetry can now be filed under the vocabulary
your organisation already uses: map a Kubernetes pod label once (tags.labels)
and it rides every signal as avuru.tag.<key>, applied at collection so
workloads nobody instrumented carry it too, then filter traces and logs by it —
a trace matches when any service that took part carries the tag. A service can
also declare its own domain, environment and avuru.tier as resource
attributes and be grouped accordingly across Kubernetes namespaces with no hub
config, with a warning when a declaration cannot be honoured rather than a
silent fallback. Two more clients read the same public API: the avuruobs
CLI, whose --fail-on predicate gates a deploy with three exit codes so a
tripped gate is distinguishable from a broken one, and a Grafana data
source — a backend plugin, so the API token never reaches a browser and
queries leave the Grafana server. Every screen now links to the page of the
manual that explains it. Inter-zone traffic accounting reports bytes per
availability-zone pair from kernel flows, standalone from the per-edge network
feature. Fixed: on a meshed cluster the map drew every application call as two
hops through a proxy, and turning on kernel network flows rendered a sensor
config the eBPF tracer refuses to parse — so the container never started.
Upgrading is a normal helm upgrade; everything new is off or unchanged by
default.
GitHub release
Open at both ends. Getting data in no longer means re-pointing every
sender you own: Jaeger (gRPC and thrift/HTTP), Zipkin, Prometheus
remote-write and Loki push receivers land beside OTLP, one values flag
each and all off by default, every one through the same tenant stage so
per-project ingest keys are enforced whatever the wire protocol. Forwarding
exporters (gateway.forward.otlp, gateway.forward.kafka) dual-write to
the backend you run today behind a bounded queue, so adopting avuru obs is a
reversible decision; every protocol is exercised in CI through the receivers a
real Helm install renders. At the other end, member projects make one
project read the union of several clusters on every screen — membership one
level deep, each viewer still seeing only the members they were granted — and
component toggles (hub.enabled, ui.enabled, gateway.enabled) let a
secondary cluster install the ingest half alone against the central store, with
impossible combinations refused at install time. Projects also gained
per-project retention (an hourly tenant-scoped trim) and per-project
storage usage in Settings → Storage. On the energy side, /green names the
nodes its numbers come from and carbon budgets report whether they can actually
reach a channel. Fixed: the hub answered with its built-in retention rather
than the configured one. Schema migration 0019 applies automatically;
upgrading is a normal helm upgrade.
GitHub release
· changelog.
Operate it from the UI. Nearly everything an operator changed through
values.yaml now lives in the app: runtime collection control (per-signal
switches the sensor follows in seconds, default-off behind a deliberately
narrow Role), service groups authored in Settings → Groups, an editable
SSO group→role mapping beside the chart's read-only rules, and personal
API tokens — hashed at rest, shown once, resolving to their owner's live
permissions so disabling a user disables every token they hold. The
Dashboard becomes the landing screen — group health, live topology, firing
alerts and Kubernetes capacity in one view — and the service map shows
real status rings from the health rollup, caller-side p50/p95 per edge, hover
focus and shareable filters. Also: sorting and filtering on the Nodes screen,
and a fix for three green endpoints that answered without authentication.
Schema migrations 0016–0018 apply automatically; upgrading is a normal
helm upgrade.
GitHub release
· changelog.
Accounts you can administer. v0.2 created users; v0.4 finishes their
lifecycle. Settings → Users edits a name and role grants, resets passwords
and deletes an account behind a disable-first rule, while a new
Settings → Account tab lets anyone rotate their own password — current
password required, other sessions evicted, yours kept alive. Password
operations are refused for SSO users, whose credential lives at the identity
provider. Reviewing that surface closed three ways into an account: a local
password that could be minted on an SSO-only user, a login lockout that
rotating IP addresses walked straight past, and an SSO login able to take over
a local account's email. Operationally, an install whose schema migration never
ran now repairs itself (hub.autoMigrate) and reports applied-versus-expected
schema in Settings → Status; a non-default ClickHouse database name works;
green survives nodes without RAPL; and login works behind a reverse proxy that
rewrites Host. Upgrading is a normal helm upgrade.
GitHub release
· changelog.
Patch. The chart's image defaults never matched what the release workflow
publishes — wrong repositories and a tag format that was never pushed — so a
helm install with no --set could not resolve the hub, UI, gateway or
TDP-estimator images. Both halves are fixed; the green estimator, which
shipped with no repository at all, now defaults like its siblings. No schema
migration, no API or configuration change.
GitHub release
· changelog.
Tenancy you can trust. v0.2 secured the read side; v0.3 closes the write
side. Projects become something you administer — create, rename and delete
them from the UI, with built-in and config-defined entries kept read-only —
and per-project ingest keys end topology-based trust: keys are validated
in the gateway, and in enforce mode the key's project is the authoritative
tenant, overriding whatever a sender claims (the default log mode changes
nothing about the pipeline, so the drop-in OTLP promise survives the upgrade).
A one-click read-only demo signs a visitor in as a scoped viewer without
the shared password ever reaching the browser. Green now works on the
RAPL-less cloud VMs most fleets run on, with every modeled number labeled
estimated end to end. The runtime collection control plane lands its
groundwork — overlay store, validated API, least-privilege RBAC — and the
deploy layer is renamed avuruops → avuruobs (breaking; see the
upgrade guide).
GitHub release
· changelog.
Depth and control. The five-minute install becomes one a real team can run every day: the hub is secure by default (login, Admin/Editor/Viewer roles granted per project, OIDC SSO with any IdP), signals are modular (one switch per family gates schema, API, pipeline, collection and UI together), and the sensor is provably safe to leave on — CI-enforced. Four new modules build on the data already collected: error tracking (deduplicated, triageable issues plus a browser ingest path), service health groups (criticality tiers with dependency propagation), alerting (SSRF-guarded webhooks on health transitions) and green energy & carbon (per-service Wh/gCO2e, budgets, a CSRD-ready export). Network health lands on the service-map edges. Licensed AGPL-3.0 from this release on. GitHub release · changelog.
The first tagged release: the wedge. A fresh Kubernetes cluster reaches a live service map in under five minutes with zero app changes — enforced as a CI gate. All four v0.1 signal tiers ship — traces (Full), logs (Basic), continuous CPU profiling (Lite, opt-in), infra metrics (Supporting) — plus the OTLP drop-in migration path, a per-project model with collection controls, and a deep trace inspect. GitHub release · changelog.
Once a version is cut, its entry links to the
GitHub release and the
matching changelog items. Releases are cut per the engine repo's
RELEASING.md.