opik-mcp

experiment-flows

Purpose

This feature lists and reads experiments and datasets, finds a case in a dataset, compares experiments case by case, and sends the writes an evaluation needs. Open it to learn what a comparison page promises, how a case is found, or what the evaluation writes send to the backend.

What it does now

Experiments

Datasets and cases

Comparing: list('dataset_item', experiment_ids=[A, B])

Evaluation writes

How it works

list/read → read_list/registry.py → entities/experiment.py
                                   → entities/dataset/__init__.py
  dataset_item, no experiment_ids → items.py (list_items, fetch_item)
  dataset_item + experiment_ids   → compare.py run_compare
      → get_experiment per id, then one gather: rows, feedback
        definitions, stats per experiment, output columns (page 1)
      → compare_guards.py (same dataset, version, status, case count)
      → figures.py (per-experiment lines), compare_notes.py, layout.py render
        (compared_row.py: one case's runs and cells)
write → writes/registry.py → writes/operations/evaluation.py build hooks

Decisions

Traps

Proven by

Log