Skip to main content

Environment

Version: 4.8 | Last updated: 2026-10-06

The Environment page answers one question: what is the state of your system right now. It reads directly from your connected observability backend, so what it shows is what Bluebox itself can see when it investigates.

Open it from Environment in the sidebar.

Needs attention​

Anything currently flagged appears at the top, one card per service or application, with the open findings behind it and the readings that led to the status. Each finding links to the finding itself, the same way it does in the table below. Each card carries an Investigate action that starts an investigation scoped to that entity.

When nothing is flagged, this section is empty: the Services count above already states how many are healthy, so there is nothing left to add here.

The Services count, here and on the Overview tile, is the number of services and applications your observability backend knew about in the last hour. Both places state that range next to the figure, and both count the same thing, so the two never disagree.

How many are flagged is a different question, and the two places answer it differently on purpose. The Overview tile counts only entities with an open finding, which is all it checks, and says so: it reads "with findings". This page goes further and also weighs each service's own error rate and each host's utilization, so it can flag a service that has no open finding against it. Expect the page to flag at least as much as the tile, never less.

Services​

The table lists every service and application in your environment, with three golden signals for each:

  • Req / min: requests per minute.
  • Error rate: the share of those requests that failed.
  • Response p95: the response time 95 percent of requests came in under.

All three cover the last hour, and each reports what its own job calls for:

  • Req / min is the typical five-minute period, so one busy stretch does not redefine what the service normally handles.
  • Error rate is the failures across the whole hour, so a problem that started twenty minutes ago is visible rather than averaged away by the calm before it.
  • Response p95 is the worst five-minute period. A percentile cannot be averaged with another percentile, so there is no p95 for the whole hour to show you, and the worst period at least describes a moment that happened.

Applications carry an Application chip so they are not read as services. Services whose own name is only a port also show the process they run on, which is often the only way to tell two of them apart.

Find a service​

Two controls, and they work together:

  • The status buttons narrow the table to one status. Each carries its own count, and only statuses present in your environment get a button, so no button ever promises rows that are not there. When everything shares one status there are no buttons at all, because there would be nothing to narrow.
  • The search box matches on name, environment, cluster, and the process a service runs on.

Applying a search while a status is selected narrows within that status rather than across the whole table.

See more about one service​

Each row carries a small chart beside its request rate, error rate and response time, showing how that signal moved across the window rather than only where it ended up.

The row names its cluster, environment and the process it runs on beneath the service name, and links its open findings. Beyond the first few, the rest are counted rather than listed, so one noisy service cannot push the others off the screen.

The actions button at the end of a row opens a menu with Investigate, which starts an investigation scoped to that service.

When a reading is missing​

A period the service reported nothing for leaves a gap in its chart rather than a line drawn straight across it, and the chart says how many gaps there were.

An error rate too small to show at one decimal place reads <0.1% rather than a dash, so a service with a failure-rate finding against it is never reported as having no errors. A dash means the window had no errors at all.

A signal that was never measured reads not measured, never 0. The distinction matters: a service with no metric series is not a service serving no traffic, and showing it as zero would state a measurement nobody took.

If a query did not complete, the page says which part is missing instead of quietly showing less. A partial read is labelled as partial.

A service your observability backend can detect but has recorded no traffic for, and that has no open findings against it, is listed with the status unknown. It is neither healthy nor in trouble: nothing has been measured for it, and saying so is the point. A service that has gone dark is often exactly what you came here to find, so the page shows it rather than leaving it out. This status only appears once the golden-signal read itself completed, so a temporary read failure never turns the whole table unknown.

Applications are judged differently, because the traffic signal does not apply to them. An application's status comes from its open findings alone, so one with nothing flagged against it reads healthy even though no request metrics were read for it.

Check which services Bluebox can fully see​

The Telemetry tile says how many of your services send all three signals Bluebox investigates with: traces, logs and metrics. It counts the same services as the Services tile, minus applications, which do not send these signals. That is why the two figures can differ.

Below the headline, each signal shows how many services send it, for example logs 9/21. When a signal is missing from some services, the tile names the one missing most often and what Bluebox cannot do without it.

A service that sends nothing at all counts as missing every signal. A signal Bluebox could not check, because the read failed or came back incomplete, shows as not measured rather than missing, and the tile counts that service as not measured instead of not covered. If Bluebox could check none of them, the tile reads not measured.

What Bluebox needs to fill this page​

Service traffic, error rate and response time are derived from distributed traces, so any service sending OpenTelemetry traces to your connected observability backend appears here. No agent installation is required.

When the page has nothing to show​

An empty page is not the same as a broken one, so the page tells you which of three situations you are in rather than showing a blank table for all of them.

  • No observability backend connected. The workspace has nothing to read from yet. The page explains what Environment shows and links you to Connections to connect one.
  • Connected, nothing reporting. The connection works and Bluebox checked for services, hosts, and databases, and found none. Nothing is wrong. The page fills in on its own once telemetry starts arriving.
  • Could not be read. A query failed, so Bluebox has no measurement either way. This says nothing about whether your telemetry is arriving. Retry runs the read again.

If your observability backend is slow to answer, the page keeps loading and retries on its own for about a minute before it reports that the read could not complete. If the page already showed your environment, it stays on screen with a note saying when it last refreshed.

  • Findings: how Bluebox surfaces issues across your environment.
  • Investigations: what happens after you select Investigate.
  • Connections: connecting Bluebox to your observability backend.