Skip to main content

API

The hub exposes a client-agnostic REST API used by the UI and any future client (CLI, Grafana data source). All responses are JSON. Times are RFC3339 (UTC); durations are milliseconds.

  • Base path — /api/v1.
  • Time range — most query endpoints accept start and end (RFC3339); omitted, they default to the last 15 minutes.
  • Projects (tenancy) — every data endpoint is project-scoped: the X-Avuru-Tenant header selects a project, omitted requests read the default project. GET /api/v1/projects lists what exists.
  • Authentication — with auth enabled (the default), endpoints require the session cookie set by POST /api/v1/auth/login or the OIDC flow. Roles (admin/editor/viewer) and per-project grants are enforced server-side; the selected project is validated against the caller's grants (403 otherwise). Read endpoints need viewer, mutations need editor, and rows marked admin need the global admin role. GET /healthz and GET /api/v1/auth/config are always open. Since v0.5.0, a personal API token sent as Authorization: Bearer avurut_… authenticates a request as its owner, with the owner's live roles and grants; the bearer path never falls through — a bad or expired token is a 401, not an anonymous session.
  • Login rate limiting — failed password attempts are counted on three axes: email + address, address, and (since v0.4.0) the account alone, so guesses spread across many addresses can no longer walk past the per-account lockout. A tripped limit returns 429 and blocks only login for that account, over a self-healing one-minute window; established sessions and successful logins are unaffected.
  • Cross-origin writes — mutating requests are checked with Origin against the Host the hub received. Behind a proxy that rewrites Host, name the real origins in auth.trustedOrigins or relax the check with auth.originCheck (enforce | log | off); an OIDC publicUrl is trusted automatically.

Endpoints​

Method & pathPurpose
GET /healthzLiveness probe (always 200, even during a ClickHouse outage).
GET /api/v1/statusHub build info + ClickHouse reachability.
GET /api/v1/capabilitiesThe install's active module set — the UI sidebar follows it.
GET /api/v1/auth/configLogin-page bootstrap: {enabled, methods, forceSSO, demoEnabled}. Always registered, auth on or off.
POST /api/v1/auth/loginLocal login — body {email, password}; sets the session cookie. Rate-limited.
POST /api/v1/auth/logoutEnd the current session (server-side, so revocation is immediate).
GET /api/v1/auth/meThe caller's identity: user (id, email, name, anonymous, origin) + grants.
POST /api/v1/auth/passwordSelf-service password change — body {currentPassword, newPassword}. The current password is required; every other session of the caller is ended and the calling session's cookie is re-minted, so you stay signed in. Local accounts only (origin=local); SSO callers get a 409 — their credential lives at the IdP. Rate-limited per account.
POST /api/v1/auth/demoOne-click read-only demo sign-in — starts a session as the demo viewer using server-held credentials (no request body). Registered only when demo mode is enabled (auth.demo.enabled); rate-limited.
GET /api/v1/auth/oidc/startBegin the OIDC authorization-code + PKCE flow — redirects to the IdP (400 if OIDC is not configured).
GET /api/v1/auth/oidc/callbackIdP redirect target — validates state, exchanges the code, starts the session.
GET /api/v1/auth/oidc/mappingThe merged OIDC group→role mapping: chart-declared rules (source: config, read-only) and UI-authored ones (source: db), each with editable and shadowed flags — a shadowed rule names a group the chart also declares; it stays visible but never grants. Registered only when OIDC is configured — no provider means nothing to map, so the honest answer is 404. Admin.
PUT /api/v1/auth/oidc/mapping/{group}Create or update an authored rule — {role, projects}. On a chart collision the rule is stored and marked shadowed rather than silently ignored or refused. Applies at the group's next sign-in or token refresh; reaches every replica within ~15 s. Admin.
DELETE /api/v1/auth/oidc/mapping/{group}Delete an authored rule (404 if none has that group). Chart-declared rules cannot be deleted here. Admin.
POST /api/v1/auth/oidc/mapping/resetDelete every authored rule, returning the mapping to exactly what the chart declares. Admin.
GET /api/v1/tokensThe caller's personal API tokens — prefix, name, createdAt, expiresAt, lastUsedAt and tokenHash (the revoke handle). Never the raw token. A global admin may list another user's with ?user=. Any signed-in user — deliberately no role floor, so a user whose grants were revoked can still clean up the credentials they handed out.
POST /api/v1/tokensMint a token — {name, expiresInDays} (0 or omitted = never expires). The 201 response is the only place the raw avurut_… value ever appears; only its SHA-256 is stored. The token authenticates as its owner's live identity on every request, so disabling the owner disables the token.
DELETE /api/v1/tokens/{hash}Revoke a token by its tokenHash. Self-service; a global admin may revoke anyone's. An unknown — or another user's — hash answers 404, not 403, which would confirm the hash exists.
GET /api/v1/usersList users with role grants and origin (local/oidc). Admin.
POST /api/v1/usersCreate a local user — {email, name, password, grants}. Admin.
PUT /api/v1/users/{id}Update a user — any of name, password, disabled, grants. A password is accepted only for local accounts; setting one on an SSO user is refused, so an admin cannot mint a credential that bypasses the identity provider. Admin.
DELETE /api/v1/users/{id}Delete a user. Permitted only once the user is disabled — a deliberate second step, so removal is never one click. For an SSO user this removes the local record only; because disabled is what the SSO callback checks, deleting a disabled SSO user undoes their lockout. Admin.
GET /api/v1/projectsProject list: default, config-defined projects (projects chart value / AVURUOBS_PROJECTS), UI-managed projects, and tenants auto-discovered from data. Each entry carries a source, a label, an editable flag, its members (a project that aggregates others) and its retentionDays (absent = it inherits the install-wide window).
POST /api/v1/projectsCreate a UI-managed project — {id, label, retentionDays?}, where id is an immutable tenant slug. A reserved (default/config) or duplicate id is a 409. Admin.
PUT /api/v1/projects/{id}Edit a UI-managed project — {label?, members?, retentionDays?}, a partial update: an omitted field keeps its stored value, so the separate editors cannot blank each other. members makes the project an aggregate that reads their union; membership is one level deep (409 otherwise). retentionDays must be 0 (inherit) or shorter than the install-wide window (400 otherwise), and is refused on an aggregate (409) — an aggregate owns no rows. default and config-defined projects are read-only (409). Admin.
DELETE /api/v1/projects/{id}Delete a UI-managed project (its telemetry ages out by retention). default and config-defined projects are read-only (409). Admin.
GET /api/v1/projects/{id}/keysList the project's ingest API keys — prefix, name, createdBy, createdAt and keyHash (the delete handle). Never returns the raw secret. Admin.
POST /api/v1/projects/{id}/keysMint an ingest API key — {name}. The 201 response is the only place the raw key ever appears; only its SHA-256 is stored. Admin.
DELETE /api/v1/projects/{id}/keys/{hash}Revoke a key by its keyHash. Takes effect within the gateway's verdict cache TTL. 204, or 404 if no live key matches. Admin.
GET /api/v1/collection/overlayThe stored collection overlay — per-signal on/off plus the shared namespace-exclusion list. Registered only when collection.runtimeControl.enabled (default off); its state is echoed in GET /api/v1/capabilities. Admin.
PUT /api/v1/collection/overlayReplace the overlay. A closed schema: anything outside it — including free-form collector config — is rejected, so the endpoint adds no injection surface. Since v0.5.0 the overlay is applied at runtime: the hub writes it through to its own sensor ConfigMaps and rolls the DaemonSet, so the sensor follows within seconds. Admin.
DELETE /api/v1/collection/overlayClear the overlay, returning collection to the chart defaults. Admin.
GET /api/v1/ratesThe one rate table: the UI-authored overlay, the chart values it sits on (served read-only — offering an edit would promise a change a helm upgrade reverts), and the merged effective result with each row marked chart or overlay. Prices both the cost and ai modules, so it is gated on neither. Admin.
PUT /api/v1/ratesReplace the overlay wholesale — there is no PATCH merge, because only a full desired state can express a deletion unambiguously. A closed schema: an unknown key is refused rather than silently stored. Compute rates move as a pair, so a table setting only CPU does not inherit a chart memory rate. Admin.
DELETE /api/v1/ratesClear the overlay, returning to whatever the chart declares. It does not un-price an estate that declared rates in values. Admin.
GET /api/v1/system/statusComponent health, per-signal storage (rows, bytes, compression, configured retention and the TTL the tables enforce) disk usage, and the ClickHouse connection (address, database, user — never the password). Includes a Schema component reporting applied versus expected migrations, so a half-migrated install names itself instead of failing every query. Also carries a project block for the project the request selects (X-Avuru-Tenant): rows, an apportioned estimatedBytes, rowsPerMinute, time bounds and the retention that applies (inherited says whether it is the project's own). An aggregate reports the union of the members the caller may see and names them. Admin.
GET /api/v1/servicesServices with RED aggregates over the window.
GET /api/v1/tagsBusiness tags (avuru.tag.*) seen in the window, each with a bounded sample of its values — filter discovery. Core: tags ride resource attributes on traces, which every install collects. Empty until an operator maps a label in tags.labels.
GET /api/v1/service-mapService nodes plus call edges (caller→callee from trace spans, enriched with OBI network flow bytes and per-edge health — RTT p95, failed connections — when infra-metrics is active). Trace-derived edges also carry client-side p50Ms/p95Ms — the caller's own view of that call, distinct from the callee's server-side p95 on the node. Both fields are omitted (not zero) on edges derived only from network flows, which have no span to measure. Nodes carry namespace (the workload's k8s.namespace.name, falling back to service.namespace) and, where they are not an ordinary application, a role: transport for a mesh proxy or gateway, virtual for a dependency derived from exit spans — a database, cache or broker that emits no telemetry of its own, further narrowed by kind (database | cache | queue) and named as a URI such as postgresql://orders-db. An application node carries neither field, so a cluster with no mesh and no derived dependency returns the shape it always did. Transport is recognised from the Kubernetes labels a mesh writes on its own data plane where they exist, and from workload names otherwise; labels only ever promote a workload to transport (a sidecar has no label of its own), and the topology config's applications list overrides both. Edges recovered across a mesh hop carry viaTransport (the proxies traversed) and collapsedCalls/collapsedErrorCount (how much of the edge came through the mesh, so a client drawing the hops themselves can subtract rather than double-count); per-edge health adds retransmits beside rttMs and failedConnections. All are omitted when absent.
GET /api/v1/tracesSearch traces (filters + keyset pagination — see below).
GET /api/v1/traces/overviewPer-(service, operation) RED metrics (count, errors, refusals, P50/P95/P99). refusedCount/refusedRate are omitted for an operation that refuses nothing.
GET /api/v1/traces/heatmapLatency × time histogram (sparse cells) over root spans.
GET /api/v1/traces/breakdownOne grouped aggregate over spans — the numbers behind the part-of-whole views. groupBy is a closed set (service, operation, kind, status, attribute:<key>, resource:<key>); anything else is a 400 naming the alternatives, never SQL built from caller text. scope picks the span population — entry (Server/Consumer: what each service was asked to serve, so it reconciles with RED), root (parentless: where traffic entered) or all — and they are not interchangeable. The response carries total over every matching span, computed before the limit, so a truncated top-N knows the size of its own tail; the tail's quantiles are absent, because a p95 cannot be recovered by subtraction. Takes the same filters as GET /api/v1/traces.
GET /api/v1/traces/{traceId}Full span tree for one trace (attributes, events, links).
GET /api/v1/traces/{traceId}/logsLogs correlated to a trace.
GET /api/v1/spans/{spanId}Resolve a span id to its containing trace (404 if unknown) — powers the span-id search.
GET /api/v1/logsSearch logs as one newest-first stream. Repeat service and workload=namespace/name to compose several subjects with OR; q, severity and tags then apply to the combined stream with AND. source may be repeated or supplied as a comma list of app, ztunnel, waypoint, and other; an unknown value is a 400. With selected subjects, the hub resolves each service to its workload and narrows proxy lines to that workload rather than widening the search. Each row keeps its emitting service and adds source (application, ztunnel, waypoint, or other). The response's resolutions say what each subject resolved to: proxiesUnavailable when no proxy source is on offer and why, proxiesMatchedBy when proxy lines were tied to the subject by name rather than by its pods (mesh-config off, cluster unread, pod list cut, no pod in the snapshot), and proxiesFallback when pagination compacted a large pod set. When nextCursor is present, resolutionToken is a short, signed, opaque token for the resolved service, pod, and waypoint branches (or the stable workload fallback used when a large pod set would exceed safe URL size); echo it as resolution with the next request so a cluster refresh cannot change the source set under the cursor. An invalid, modified, or expired token returns 400. A single service without source keeps the original exact-service behavior. limit, opaque cursor, start and end apply as usual.
GET /api/v1/logs/servicesSuggestions for the multi-service log selector in the requested start/end window. services includes names present only in logs. workloads contains sorted namespace/name values only when a log-backed service in the authorized project resolves to that workload; resolution uses telemetry attributes first and may use the readable cluster snapshot when mesh-config is active. The cluster snapshot alone never adds a suggestion. Returns two sorted arrays. Module logs.
GET /api/v1/services/{service}/logsOne service's logs, composed the way GET /api/v1/mesh/workloads/{namespace}/{name}/logs composes a workload's, for the common case where a service's pod files its lines under a different name than its spans carry (orders-service running as the Deployment orders). The hub ties the service to a workload: first from the service's own spans (k8s.deployment.name, then k8s.statefulset.name, then k8s.daemonset.name, weighted by span count, with k8s.namespace.name); then, with mesh-config, from the cluster snapshot: a Kubernetes Service of that name and the first workload its selector matches, or else a workload of that name. The namespace from the spans narrows the lookup when there is one; otherwise a name that exists in two namespaces is not a match. The app source is then the union of the service name and the workload's name and name.namespace, because the app's lines may be filed under either. Same params and page shape as the workload route (start/end, limit default 100, cursor, q, severity, source with anything unknown a 400, waypoint), plus workload and namespace: sources in the response carries the resolved pair, and a client echoes it back on later pages so the hub does not resolve again and the source set cannot change under a live cursor. When no workload resolves, only the service's own lines are read and sources.proxiesUnavailable says why. Module logs only; the proxy sources need mesh-config.
GET /api/v1/metrics/redBucketed RED series per service — service (empty = busiest top), points, includeAux.
GET /api/v1/health/groupsConsolidated group health over the window: per-service status from RED, tier-grouped, with critical-dependency propagation. Module service-health. includeAux.
GET /api/v1/health/groups/{name}One group's health, including each member's base vs. effective status and its critical-dependency chain (404 if the group is absent).
GET /api/v1/checksThe endpoint checks declared on the service-health config: id, the group and tier they answer for, url, and the resolved interval (the default is reported explicitly, never left blank). Module service-health.
GET /api/v1/checks/{id}/resultsOne check's recent probe outcomes, newest first: ok, status, latencyMs, error, and the traceId of the span the probe emitted for its own request — the click-through from a failed check to the request that failed. limit (default 100, capped at 500). An id that was never declared is a 404, not an empty list: "has not run yet" and "no such check" send you to different places.
GET /api/v1/auth/permissionsThe role model and, per area of the product, the lowest role that can read it and the lowest that can change it — derived by the hub from the guards its routes registered with. Any signed-in caller.
GET /api/v1/service-groupsThe group definitions (name, tier, selector) as opposed to their health: chart-declared groups tagged source: config and read-only, UI-authored ones source: db. Module service-health.
POST /api/v1/service-groupsCreate a group. Admin role. 409 when the name is chart-declared or already taken.
PUT /api/v1/service-groups/{name}Edit a group's tier and selector. Admin role. 409 on a chart-declared name, 404 if no authored group has it.
DELETE /api/v1/service-groups/{name}Delete a group. Admin role. Same 409/404 rules as the update.
GET /api/v1/alertsCurrently-firing alerts + recent fire/resolve history. Module alerting.
GET /api/v1/alerts/rulesThe configured alerting rules and channels (channel secrets redacted). Module alerting.
GET /api/v1/alerts/channelsThe configured notification channels (secrets redacted). Module alerting.
POST /api/v1/alerts/channelsCreate a notification channel. Admin role.
PUT /api/v1/alerts/channels/{name}Update a notification channel. Admin role.
DELETE /api/v1/alerts/channels/{name}Delete a notification channel. Admin role.
POST /api/v1/alerts/channels/{name}/testSend a test notification through a channel (same SSRF policy as the evaluator; 503 if delivery is not configured). Admin role.
GET /api/v1/errors/issuesDeduplicated error issues — status, service, q, sort filters, limit (default 100, max 500). Module error-tracking.
GET /api/v1/errors/statsWhat the filtered issue set amounts to: matching issues, how many are new in the window (first seen inside it) and how many regressed, total occurrences with a histogram (points, default 48), and the busiest services (top, default 5). Takes the same status/service/q filters as the list and aggregates the same issue set through the same query, so the two can never state different totals.
GET /api/v1/errors/issues/{fingerprint}One issue: stack trace, first/last seen, count, source, representative trace.
GET /api/v1/errors/issues/{fingerprint}/eventsOccurrences of an issue (keyset paginated).
GET /api/v1/errors/issues/{fingerprint}/histogramOccurrence histogram over the window.
POST /api/v1/errors/issues/{fingerprint}/statusTriage an issue — unresolved / resolved / ignored.
GET /api/v1/green/summaryPer-service energy (Wh) and carbon (gCO2e) over the window, with the coverage ratio and unattributed bucket. Coverage also carries one row per known node — its Wh and the measured/estimated split — with nodes reporting nothing listed at 0, not omitted. Module green (requires infra-metrics).
GET /api/v1/green/budgetsMonthly carbon budgets per service group: used, projected, ratio, and warn/exceeded status. Each budget also reports its deliverability (wired, alerting off, no channel, unknown channel), and the response carries warnings for budgets targeting a group nothing rolls up to — those evaluate forever at zero. Module green.
GET /api/v1/green/reportCSRD-ready export — per-app numbers plus a methodology block (formula, factor provenance, coverage). Module green.
GET /api/v1/cost/workloadsPer workload, the CPU and memory it reserved against what it used, ranked by the gap. idleCpuCores/idleMemBytes are reserved minus the peak (never the mean — a request cannot be cut below what the workload reached), and never negative. requestsNothing marks a workload that declared no request at all, which is a different problem from a small one. priced leads the response: money fields are present only when both cost.rates are configured, because a zero under a currency header reads as free. Module cost (requires infra-metrics).
GET /api/v1/cost/nodesPer node, allocatable capacity against the share claimed by container requests and the share actually in use — a node can be fully requested and almost idle, and only one of those stops it taking pods. Module cost (requires infra-metrics).
GET /api/v1/ai/summaryThe window's model calls as one line: calls, input/output tokens, failure and truncation rates, distinct models, and the coverage the totals exclude. totalCost is the sum of the priced models only and unpricedModels names what that floor is missing. Module ai.
GET /api/v1/ai/modelsPer model: provider, calls, tokens, p50/p95/p99, failed, refused and truncated — a successful call cut off at the token ceiling, reported apart from failure because it is a different thing to fix. callsWithoutUsage is the population excluded from the token sums (a call that reported nothing is not a call that used nothing), callsFromRequestModel marks rows attributed to the model that was asked for because nothing said what answered, and callsWithContent counts calls still carrying prompt or completion text — reported, never rendered. cost is absent rather than zero when no rate is declared, and pricedByPrefix marks a cost derived from a prefix rule instead of an exact entry. Module ai.
GET /api/v1/ai/callersPer (service, model): calls, tokens, failures, truncation and cost — spend with an owner. A row whose every call reported no usage carries no token total and no cost, because a zero in either would read as free. Module ai.
GET /api/v1/ai/toolsPer tool executed inside an agent turn (gen_ai.tool.name, falling back to the span name): calls, failures, refusals, p50/p95/p99, and the services invoking it. A tool a turn hit four times is one row with a count — the loop is the thing worth seeing. No tokens and no cost: a tool execution spends neither, and a zero would read as free rather than as not the unit. namedBySpan counts calls named from the span rather than the attribute — a weaker attribution, stated. modelFilterIgnored reports that a model filter was set and could not apply, because a tool span carries no model. Module ai.
POST /mcpThe Model Context Protocol endpoint — not under /api/v1, because MCP clients are configured with a server URL and the protocol owns the path. Six read-only tools (service_context, search_traces, get_trace, search_logs, list_error_issues, list_services) over data the install already stores; a tool whose module is off is absent from the tool list rather than present and failing. A service argument is resolved against every signal, not only entry spans, so a workload that ships logs and no traces can be asked about by name; an unknown name is still an error naming the near matches rather than an empty page. service_context then omits red rather than reporting zeros it cannot compute, and reports what the service did send under signals, while search_traces returns a note saying the service is untraced — an empty trace list must never read as an outage. Authenticated with a personal API token, resolving its owner's live permissions and project grants — the same reads that person has in the UI. Every call is logged with the token owner, the tool, its arguments and the row count, never the content returned. Module mcp, off by default.
GET /.well-known/oauth-protected-resource
GET /.well-known/oauth-protected-resource/mcp
GET /.well-known/oauth-authorization-server
OAuth 2.1 discovery, unauthenticated and at the origin root because a client fetches them before it has any credential; the /mcp-suffixed form is the one RFC 9728 has a client derive from the resource's own path, and both answer the same document. On a Helm install the Ingress and the UI's nginx route all of them to the hub alongside /mcp, since no /api rule can cover a path fixed at the root. Served only when modules.mcp.oauth.enabled is on — advertising a flow an install does not run would send every connector down a dead end. code_challenge_methods_supported is exactly ["S256"].
POST /api/v1/auth/oauth/registerRFC 7591 dynamic client registration. Open by necessity — a hosted client has no credential for your install and no way to be pre-registered — and it grants nothing: the registration can read no data at all until a person consents to it by name. Bounded by rate limit, redirect-URI count and body size.
GET /api/v1/auth/oauth/authorize
POST /api/v1/auth/oauth/token
Authorization code with PKCE (S256 required), and the exchange. A code is single-use and bound to its client, its exact redirect URI, its challenge and its resource; presenting one twice revokes the whole grant. Access tokens are opaque and audience-bound to POST /mcp — one is refused on every other route.
POST /api/v1/auth/oauth/revokeRFC 7009 revocation, for a client that disconnects itself. Presenting either half of a token pair revokes every token issued under that grant — a client disconnecting means the pair, not one of them. The answer is 200 whether or not the token existed, as the RFC requires, so the endpoint confirms nothing about what is live; Cache-Control: no-store on every response.
GET /api/v1/auth/oauth/grants
DELETE /api/v1/auth/oauth/grants/{id}
The applications the caller has connected, and disconnecting one. Revokes the consent and every token issued under it, so the application stops reading on its next request.
GET /api/v1/infra/nodesNode utilization (kubeletstats): latest CPU/memory/network, pod counts, points-bucketed series.
GET /api/v1/infra/podsPod utilization, busiest first — node scopes to one node, limit caps rows.
GET /api/v1/infra/pod-connectionsTrace-backed Pod connections. Params: start, end, optional node (either endpoint), limit (1–500, default 200). Returns connections[], limit, truncated. Each connection carries source/target (name, namespace, node, service), optional unresolved peer, caller-side calls, errors, p95Ms. Matches Client/Server identities within the same tenant; excludes auxiliary calls and deduplicates caller spans. An unresolved peer does not assert public Internet location. Requires viewer access to the project and the infra-metrics module.
GET /api/v1/agentsSensor inventory: per-node telemetry freshness per signal — windowSec (default 600) bounds "fresh".
GET /api/v1/mesh/proxiesThe mesh's own workloads — those classified as transport — with their RED and the call volume they carry, callsIn and callsOut counted apart: traffic arriving with none leaving is a proxy that has stopped forwarding, which its own error rate need not show. No new query; this is the service list with the topology classifier's verdict inverted. Each proxy also carries a role (control-plane, ingress-gateway, egress-gateway, gateway, waypoint, ztunnel, sidecar) and the namespace it runs in — both omitted rather than guessed when the workload's labels and name settle neither. Roles come from the same avuru.transport.* labels that classify transport for the map, with names as the fallback; gateway is its own value because the Gateway API label alone establishes that a workload is a gateway and not which direction it faces. With the infra-metrics module active, bytesIn/bytesOut report what the proxy actually moved (kernel flows, not spans) and rttMs, failedConnections and retransmits the health of the links it moved them over — rttMs being the worst link rather than an average, the failures summed. All five are omitted, never zero, when nothing measured them. Module mesh.
GET /api/v1/mesh/control-planeControl-plane health over the window: connectedProxies, pushes, rejectedConfigs and convergenceP95Ms. available leads the response and every other field is omitted when it is false, with a reason instead — a control plane nobody scrapes would otherwise report zero rejected configs, which reads as perfect health. state says why it is silent — unconfigured (nothing is scraping), unreachable (scraped, no answer) or unrecognised (answered, none of the metrics this product reads came back) — and kind names the control plane whose metrics were understood (istio), absent when none were. pushP95Ms reports how long a push itself takes, which convergence alone conflates with how long proxies take to apply it; writeTimeouts counts the pushes that never landed because a proxy was too slow to receive them; and configEvents is the configuration churn the control plane is absorbing. Each is omitted when its series is not in the keep-list or not in the window, so an older scrape configuration reports fewer fields rather than zeros. Module mesh; the numbers additionally need infra-metrics (the scrape lands in the metrics tables) and mesh.controlPlane.enabled.
GET /api/v1/mesh/namespacesEvery namespace the cluster defines, not merely those that sent telemetry — which is the point: a namespace enrolled in the mesh with nothing behind it emits nothing, and that silence is a misconfiguration no other endpoint here can report. Each row carries its dataplaneMode (ambient | sidecar | none, read from the namespace's own labels, with ambient winning during a migration), its waypoint, its mtlsMode with mtlsSource (mesh, namespace) and the mtlsPolicy that decided it, workloads and enrolled counts from pods (absent when pods were not read), the services and workloads telemetry saw in the window, and its errors/warnings finding counts. state leads the response and says why a read failed — ok, unconfigured (module on, hub outside a cluster), forbidden (the ClusterRole was not granted, naming it) or no-crds (the mesh's CRDs are not installed) — with reason in the operator's terms. Telemetry is joined on (namespace) by the same resolution the service map uses, so the two can never disagree. Module mesh-config.
GET /api/v1/mesh/configThe Istio and Gateway API objects read from the cluster, with the findings against each. Without name it lists objects — kind, namespace, name, errors/warnings — filterable by kind and namespace; with kind, namespace and name it returns that one object's spec alongside its findings, each a {code, severity, message, hint, ref}: what is wrong, what to do, and which object to open. Codes are this product's own — fifteen configuration codes, from MESH_ROUTE_BACKEND_MISSING to MESH_PRINCIPAL_UNKNOWN, listed with what each catches on the mesh page; each finding's ref carries its kind, so a Service and a Workload spelled the same way open on the right tab. The five checks that read pods go silent, and checksSkipped says so, when pods were not readable or the pod list was cut. Objects are served from an informer cache — no table, no retention — with managedFields and the last-applied annotation stripped on ingest, and a truncated flag rather than a silently short list on a cluster past the snapshot cap. Module mesh-config.
GET /api/v1/mesh/securityWhat the cluster declares about mutual TLS beside what the proxies observed, per workload and per edge, over the window (includeAux). available and state lead — unconfigured (nothing scrapes the data plane), unreachable (proxies not answering, with targets.down naming the pods) or unrecognised — because a data plane nobody scrapes reports zero plaintext, which reads as a fully encrypted mesh; declared says whether the configuration half was read at all. Each workloads[] row carries observed (mtls, plaintext, unknown units, mtlsShare absent when nothing was classified — 0/0 is not 0 %), declaredMode with its declaredScope, a posture (strict-and-mtls, declared-strict-observed-plaintext, permissive-but-all-mtls, permissive-with-plaintext-callers, uncarried, observed-only-*, idle, unknown), its plaintextCallers and its findings (MESH_MTLS_NOT_ENFORCED, MESH_MTLS_READY_TO_TIGHTEN, MESH_PLAINTEXT_CALLERS, MESH_TRAFFIC_UNCARRIED); edges[] carries the same counts per caller→callee as the destination's proxy reported them, and findings is the flat list. The join is one fold in the API layer, read by every screen. Module mesh with infra-metrics; the declared half needs mesh-config.
GET /api/v1/mesh/workloadsEvery workload the cluster runs, from configuration, joined to whatever telemetry saw — namespace scopes, mode filters (ambient, sidecar, none, or declared-only: asked for by its labels, running, and neither injected nor captured — the enrolment gap by itself). Each row carries declaredMode beside dataplaneMode, injected and captured as pods report them, waypoint with waypointSource, pods/runningPods, declaredMtls (mode, source, policy — absent when no policy applies), observedMtls (absent when nothing measured it: not measured is not zero), posture from the same fold as the security route, hasTraffic with ratePerSec/errorRate as pointers, its services, its covering policies and its errors/warnings. state/reason lead as on every mesh-config response; podsTruncated and checksSkipped say, in one sentence, which pod-gated checks did not run and why. Module mesh-config.
GET /api/v1/mesh/workloads/{namespace}/{name}One workload whole: the row above — now with createdAt and createdFrom (controller, or pods when the controller was not read and the oldest pod dates it), app and version from the labels every mesh tool reads — its own findings, the covering policies' findings beside them, and its pods (name, node, phase, revision — the template hash, i.e. the rollout — createdAt, injected, captured) bounded to 50 with podsShown/podsTotal. The page adds the cluster's record: labels (without the template hash), annotations from the controller within the hub's bounds (64 keys, 2 KB per value, annotationsCut when they cut), a health verdict (healthy | degraded | down | idle in the health board's vocabulary, with the reason that decided it), and routes: every HTTPRoute, GRPCRoute, VirtualService and DestinationRule that reaches the workload through one of its Services, each with the service and host it came through and its own findings. workload is absent when the cluster or its pods could not be read — a zero-valued workload would read as one that runs nothing. Module mesh-config.
GET /api/v1/mesh/workloads/{namespace}/{name}/logsOne workload's logs from three sources as one stream: its own lines (the service named as the sensor names it, name or name.namespace), the ztunnel lines that name one of its pods, and the waypoint lines that name its Service (name.namespace.svc) — one query with a branch per source, so the page shape and the keyset cursor are the logs route's own (logs[], nextCursor). Params: start/end, limit (default 100), cursor, q (case-insensitive substring), severity (minimum), source (comma list of app, ztunnel, waypoint; default all; anything else is a 400) and waypoint=namespace/name for an install that cannot resolve the binding itself. sources says what was actually asked: the service names per source, the needles a proxy line had to contain, precise when the pods came from the mesh-config snapshot and, when they could not — module off, cluster unread, pods unreadable or cut, workload absent — a fallback sentence, with the proxies' lines then matched by name and namespace together so a workload stays apart from its namesake elsewhere. Modules mesh and logs; mesh-config only makes it precise.
GET /api/v1/mesh/workloads/{namespace}/{name}/requestsOne workload's requests as its proxy counted them, over the window: responseFlags (flag, requests, meaning in words — UO circuit breaker open, URX retries exhausted, UF, UT, NR, DC, RL; an unknown flag passes through verbatim), destinationVersions, and callers with errors5xx. measured says whether this workload had series — whether the data plane is read at all is the security route's to say — and upstreamStatsHint is always present: per-upstream counters (pending overflow, outlier ejections) are not collected, and it names the proxy setting that exposes them. Module mesh with infra-metrics.
GET /api/v1/mesh/waypoints/{namespace}/{name}What a waypoint serves, from the inventory's bindings — namespaces, services and workloads bound to it — with its scope (absent when no Gateway of that name was read: the MESH_WAYPOINT_MISSING case) and running, a pointer because not running and could not look at pods are different instructions. The question a waypoint's own traffic cannot answer: one nothing is bound to and one whose clients are idle look the same on the wire. Module mesh-config.
GET /api/v1/network/zonesBytes exchanged per availability-zone pair over the window, from the sensor's inter-zone counters. Not part of /service-map: zones are node topology, not graph elements. Requires the infra-metrics module and sensor.obi.network.interZone.enabled.
GET /api/v1/profiles/servicesServices with CPU profiling samples in the window, busiest first.
GET /api/v1/profiles/flamegraphOne service's aggregated flame graph (service required) — inclusive values, self at leaves.
POST /v1development/profilesOTLP profiles ingest (protobuf) — the sensor's profiler exports here. Deliberately outside /api/v1: it is the otlphttp exporter's default profiles path, and a documented profiles-only exception to "the hub is never in the telemetry byte-path" until the ClickHouse exporter supports profiles.
POST /internal/v1/ingest-keys/validateNot a client API. The gateway's ingest-key validation call, guarded by a shared token the chart generates; it is registered only when that token is configured. Listed here so the route is not mistaken for an unauthenticated hole: it answers {valid, project} for a key and is the reason the hub stays out of the telemetry byte-path.

GET /api/v1/traces accepts:

ParamMeaning
service, operationFilter by service / operation. service matches traces the service participates in, not only the ones it roots.
statusok, refused or error. refused is a server-side 4xx — a request the server turned away (a WAF block, an authorization denial). It is counted apart from errors, never inside them, so it never moves the error rate RED and alerting are built on; a client-side 4xx is an error, as the HTTP semantic conventions have it. ok means neither errored nor refused.
tagsAttribute equality, comma-separated: http.status_code=500,http.method=GET. Keys under the reserved business-tag prefix (avuru.tag.*) are matched against the emitting workload's resource attributes instead, and on traces they match by participation — the trace is returned when any service that took part carries the tag, the same rule the service filter uses. One filter string therefore means the same thing on /traces and /logs.
ordernewest (default), oldest, or slowest.
minDurationMs, maxDurationMsDuration band.
includeAuxtrue to include auxiliary traffic (health checks, /actuator/*, metrics). Excluded by default.
limit, cursorPage size and opaque keyset cursor (nextCursor in the response).

GET /api/v1/traces/heatmap additionally takes tags, includeAux, timeBuckets and durationBuckets. GET /api/v1/logs takes repeated service and workload, plus source, severity, q (full-text), tags, limit and cursor. See the endpoint table above for composition and source resolution rules.

:::note Planned Remote configuration (OpAMP) and streaming (WebSocket) surfaces are on the roadmap and will be documented as their UIs arrive. :::