Prove that each tool returns what the user asked for once a real Opik answers. The unit, conformance and stdio e2e suites run against stubs; they prove the wire contract and cannot see a filter that selects the wrong traces or a read too large for the host to accept. When a user-flow evaluation fails, this suite says whether the tool or the model was wrong.
scripts/seed_e2e_backend.py writes a fixed fixture to any Opik over its
REST API, verifies the backend holds it, and returns a manifest of every
name, id and count a test may assert. It imports nothing from opik_mcp, so
a bug in the server cannot corrupt the data that tests it.tests/live/ drives python -m opik_mcp over stdio with a real MCP client
and asserts exact values from the manifest. Marker live, run by make live..github/workflows/live.yaml runs it on every pull request, on main and
nightly, and posts failures on main, nightly and dispatched runs to Slack.scripts/live_alert.py builds that Slack message. It leads with what failed
on which backend and the likely cause, then the buttons to the run, each
failing job’s log and the commits since the last green run, then each
failing test with its file and assertion and the next steps, including a
command that reruns exactly those tests.A fresh fixture per run. Every run seeds its own fixture: one project named
after the run, plus datasets, experiments, prompts and rules carrying the same
name, all under mcp-live-run-<run>. It takes seconds locally and under a
minute on cloud, and the run deletes it at the end with everything its writes
created. No fixture outlives a run, so none can go stale or drift in a shared
workspace. The online rules are created disabled, so cloud never runs the
seeded LLM judges. The fixture has a content axis (errored traces, threads,
scores, two experiments with known regressions, Diagnostics issues) and a size
axis: a tiny, a typical, a heavy and a wide trace, a short and a long thread, a
small and a wide dataset, a prompt with a long history, more rules and score
names than the project summary lists. Each large case crosses the limit a read
has for it.
Ids carry time. since and until filter on the time inside a record’s
UUIDv7 id, not on start_time. The seed mints each id from its record’s own
instant across two windows of one length: 7 days against a local backend,
10 hours against any other, because Opik cloud refuses an id more than about a
day old. The tests ask for the manifest’s exact instants, so the window tests
run on both backends.
Tests go in through the host’s door only. Nothing under tests/live/
imports opik_mcp. The suite depends on the tool names, their arguments and
the answer shapes, which the conformance snapshots pin, so moving modules
inside src/ does not touch it.
Writes never touch the fixture. Every record a write test creates,
including the thread it closes and the issue it resolves, lives in a sibling
project of the fixture, mcp-live-run-<run>-w, so no write moves a count a
read asserts. Before it starts, a run sweeps what a crashed run left behind
more than two hours ago. Runs do not use the e2e-cuj- prefix, which the
shared cloud workspace’s own cleanup sweeps on its own schedule and could take
mid-run.
Sizes. Every answer’s size goes to the job summary. The size tests assert
that each answer stays under what Claude Code accepts, the ceiling defined in
tests/live/conftest.py, and that every cut states a count and the call that
gets the rest.
Two jobs. live-local starts an open source Opik at its latest release
from GHCR images with opik.sh --backend --port-mapping, seeds it and runs the
suite. live-prod runs the same suite against a shared cloud workspace and
prints the cloud version next to the open source one. It skips with a notice
until OPIK_E2E_API_KEY and OPIK_E2E_WORKSPACE exist.
Start an Opik backend, point OPIK_URL at it, run make live:
cd <opik checkout> && OPIK_VERSION=<release> ./opik.sh --backend --port-mapping
OPIK_URL=http://localhost:8080 make live
On a machine that already runs other Opik stacks, use an isolated compose
project that publishes only the backend on loopback. On Apple Silicon the
minio image’s arm64 build has no wget, so its healthcheck never passes; the
amd64 build under emulation does:
# live.override.yaml, next to docker-compose.yaml
services:
backend:
ports: !override
- "127.0.0.1:28080:8080"
minio:
platform: linux/amd64
cd <opik checkout>/deployment/docker-compose
OPIK_VERSION=<release> docker compose -p opik-mcp-live \
-f docker-compose.yaml -f live.override.yaml --profile backend up -d
OPIK_URL=http://127.0.0.1:28080 make live
To seed a fixture by hand, to explore it or for another harness to reuse, and to delete it again:
OPIK_URL=http://127.0.0.1:28080 uv run python scripts/seed_e2e_backend.py --prefix my-fixture
OPIK_URL=http://127.0.0.1:28080 uv run python scripts/seed_e2e_backend.py --prefix my-fixture --wipe
list('project_metric') and where the backend cuts a span body. The
comparison’s constant cost on a 100,000-case suite stays hermetic: a run
cannot seed a suite that size, and the claim is about our page, not the
backend’s data.tests/live/test_traces.py,
tests/live/test_threads.py, tests/live/test_project.py,
tests/live/test_evaluation.py.tests/live/test_writes.py.tests/live/test_sizes.py, including
test_an_answer_fits_what_the_host_accepts,
test_a_wide_trace_accounts_for_every_span and
test_the_backend_cuts_a_span_body_where_the_read_counts_it.test_an_error_rate_bucket_is_the_share_of_its_traces_that_errored
and test_a_sub_cent_cost_survives_the_table in tests/live/test_project.py.tests/repo/test_seed_e2e_backend.py.tests/repo/test_live_alert.py.