Optimizer Benchmark¶
Races the registered optimizers against a fast, deterministic stand-in for a real RAG optimization run — no GPUs, no LLM judge, seconds per iteration. A real 50-eval optimizer run takes hours because every eval pays remote LLM-judge calls per sub-metric; the benchmark makes head-to-head optimizer comparisons across many seeds practical.
- Optimizer types (
OPTIMIZER_REGISTRYin optimizer/__init__.py):sobol,greedy,ax_qnehvi,ax_qnehvi_mixedgp,smac_mobo,agent_qnehvi,agent_qnehvi_slo,greedy_then_hv_anchor_agent_qnehvi(what each one is: config_schema.md §4). - Benchmark function: exactly one is registered —
rag_stack_historical_surrogate(synthetic_functions/rag_stack_historical_surrogate.py). The earlier analytical testbeds (dtlz2, zdt*, branin_currin, penicillin, nasbench201, synthetic_rag, yahpo_gym, …) were removed: synthetic_rag was shown not to transfer to real e2e behavior, and the rest were generic MOBO benchmarks unused by the RAG optimizer work.
The benchmark function — rag_stack_historical_surrogate¶
Exposes the real RAG search space (parsed from a rag-stack YAML config via
the same RagSearchSpace / PerformanceContext the live Controller uses) and
scores each suggested config with:
- Quality — GP surrogate on real historical data. One independent GP per
sub-metric (
MixedSingleTaskGPwhen the space has unordered categorical dims), trained on(config, sub-metric)rows pooled from finished optimizer runs'all_evaluations.csv. The default pool is the four runs under benchmarks/historical_surrogate/ (≈195 rows; one shared corpus, judge, and search space, so pooling is a clean coverage gain). Each query returns the posterior mean (so the function is deterministic by default), then folds the sub-metrics into the scalarcombinedvia the sameaggregate_quality_metricsthe live evaluator uses. - Performance — no surrogate. Calls the real cost model
(
RAGCMAssembly.evaluate_static) on the resolved pipeline config — the exact live optimizer path, fully deterministic and already calibrated to the YAML's hardware.
Because quality matches the live trajectory within ~LLM-judge noise and performance is bit-identical, hypervolume computed on the benchmark is directly comparable to hypervolume from a full live run on the same YAML.
Constructor params accepted in the sweep YAML:
| Param | Default | Meaning |
|---|---|---|
yaml_path |
required | rag-stack YAML — defines search space, cost-model hardware, and the combined_quality aggregation rule |
historical_eval_dirs |
4-run pool | project dirs whose all_evaluations.csv to pool; shared data/qa.parquet + data/corpus.parquet must sit next to the run dirs |
sub_metrics |
7-metric set (deepeval × 5 + retrieval recall/ndcg) | sub-metric columns the surrogate models |
stochastic / noise_scale / rng_seed |
false / 1.0 / — |
inject GP-posterior-σ-scaled noise per evaluate (reproducible per rng_seed) so optimizers can be discriminated on noisy landscapes |
surrogate_difficulty |
0.0 |
0–1 blend weight of a per-metric GBM layer that restores the local ruggedness (categorical phase transitions, threshold effects) the GP smooths away |
use_gpu |
false |
fit/query the quality GPs on CUDA |
The GPs are refit on construction (~1–2 s) — nothing is cached to disk.
Matrix-sweep YAML (optimizers × function variants × seeds)¶
One config runs the full cartesian product, writes one CSV per combination,
and (unless opted out via defaults.auto_analyze: false) auto-generates the
comparison plots:
python -m rag_stack benchmark \
--config configs/optimizer_benchmark/sweep_agent_historical_surrogate.yaml
# (--plot exists but only affects the legacy single-run config shape;
# matrix sweeps plot via defaults.auto_analyze)
Schema (abridged from sweep_agent_historical_surrogate.yaml):
defaults:
gt_budget: 50 # total GT evals per (optimizer, seed)
checkpoint_interval: 5
output_dir: tmp_outputs/optimizer_benchmark_results/agent_historical
seeds: [42, 43, 44]
auto_analyze: true # default true; set false to keep CSV-only
optimizers:
- type: ax_qnehvi_mixedgp
name: ax_mixedgp_golden # REQUIRED per-experiment id — CSV prefix +
params: {n_both_init: 10} # experiments.yaml key; enables same-type ablations
- type: agent_qnehvi
name: agent_mo_inject
params:
n_both_init: 10
# The multi-objective Agent and candidate injection are intrinsic to this type.
gt_budget: 30 # per-optimizer override
synthetic_functions:
- type: rag_stack_historical_surrogate
name: rag_historical_surrogate # = subfolder + pareto-cache key
params:
# yaml_path must define a search space the pooled historical CSV rows
# decode under (rows that fail to encode are silently skipped)
yaml_path: configs/rag_stack/config_reference.yaml
# historical_eval_dirs omitted → default 4-run pool under
# benchmarks/historical_surrogate/
sub_metrics: [deepeval_answer_correctness, deepeval_answer_relevancy]
# Optional: drop specific (optimizer, function) combinations
# skip:
# - optimizer: ax_qnehvi_mixedgp
# synthetic_function: rag_historical_surrogate
synthetic_functions stays a list even though only one function class
exists — different params variants of the surrogate (e.g. another
yaml_path, a surrogate_difficulty ablation) are separate entries with
distinct names, and each gets its own subfolder + Pareto cache.
A legacy single-run shape is also accepted: one optimizer: block + one
synthetic_function: block + top-level gt_budget / seed.
Key points:
- The Agent may use its first response to request targeted evidence from the complete evaluation ledger before proposing candidates. This read-only follow-up is always available, limited to one additional LLM request, and has no configuration switch; qLogNEHVI still decides which candidate is evaluated.
- Every post-initialization optimization iteration consults the Agent exactly once. A stagnation-triggered forced iteration uses the specialized maximum-gap Agent call instead of an additional ordinary call; there is no cadence or warmup parameter.
- The single forced safety channel becomes due after three consecutive ordinary evaluations add no realized normalized hypervolume. It freezes that episode's objective normalization, joins the gap-directed Agent proposals with the same round's DOE OAT/pair cells, and selects one legal candidate by exact posterior expected coverage of the largest empty Pareto rectangle. Sources have no quota or weight. A valid forced evaluation resets the episode; an invalid or unscored attempt is blacklisted without consuming it.
- Each run produces
{optimizer_name}__{function_name}.csvunder<output_dir>/<function_name>/— the seed is a column inside the CSV (all seeds of one experiment accumulate in the same file), not part of the filename.optimizer_nameis the requirednamefield. - Per-optimizer override keys (
gt_budget/checkpoint_interval/output_dir) win overdefaults. - Seeds live only in
defaults.seeds— the matrix stays rectangular. - Failures are isolated per run: a failing optimizer logs a traceback and the sweep continues; final summary reports
N/N succeeded. - Validation is fail-fast: unknown
type, duplicatename, missingname, or unrecognized keys at the optimizer level raise before any run starts.
Re-plot only (CSVs already exist)¶
Reference Pareto cache¶
The historical surrogate has no analytical true Pareto, so the reference used
to compute hv_diff and igd (regret metrics, lower = better) is stored
per-benchmark-name alongside the run outputs — seeded with NSGA-II on the
very first build, then refreshed from the pooled evaluation trails:
<output_dir>/<name>/
experiments.yaml ← function config snapshot + per-optimizer entries
_pareto_cache.npy ← frozen reference (n_pareto × n_obj)
<optimizer>__<name>.csv ← per-iteration snapshot (seed = a column)
<optimizer>__<name>__evals.csv ← per-eval ground-truth trail
<name> is the YAML's synthetic_functions[].name field — different params
under the same function class get different names and therefore separate
caches. BenchmarkRunner lazily loads the frozen cache (falling back to
compute_true_pareto on a miss); at run time the snapshot CSVs record the
live hv only — hv_diff / igd stay nan until the backfill below fills
them against the cached reference.
Initial build (brand-new benchmark, no eval CSVs yet)¶
Use --no-pool for the very first build before any sweep has run — there's
nothing to pool from, so the reference is computed with pymoo NSGA-II on the
surrogate (~1–2 min per function):
--name is repeatable to build several benchmarks in one call.
Refresh after running new benchmark sweeps (default mode)¶
After a sweep finishes (new optimizers / seeds / runs producing fresh
*__<name>__evals.csv files in the subfolder), the default invocation
rebuilds the reference from those evals and auto-backfills hv_diff /
igd in every snapshot CSV so historical chart numbers match the new cache:
The eval CSVs to pool are read from <output_dir>/<name>/*__<name>__evals.csv
— --name and --output-dir together fully locate them, no separate
--evals-dir to set.
The default output dir comes from .rag_stack_settings.json's
optimizer_benchmark_dir field (see
System settings) (or
tmp_outputs/optimizer_benchmark_results when empty / file missing) and is
shared by all three subcommands:
// .rag_stack_settings.json — sticky default, applies to all benchmark subcommands
{
"optimizer_benchmark_dir": "/my/results/root",
...
}
To override per call, pass --output-dir /my/results/root to either
pareto-cache or recompute-metrics.
End to end, per --name:
- Reads
<output_dir>/<name>/experiments.yamlto get the functiontype+params, instantiates it. - Pools every
*__<name>__evals.csvin<output_dir>/<name>/, non-dominates the pooled points, and writes them as the new<output_dir>/<name>/_pareto_cache.npy— in this default mode the pool replaces the previous reference; NSGA-II runs only under--no-pool. - Calls
recompute_metrics_from_evalson the same subfolder to backfill historical snapshot CSVs in place (originals saved to<file>.bak_pre_recompute).
Backfill only (cache already correct)¶
If the cache is fine but historical snapshot CSVs are stale relative to it
(e.g. after manually editing the .npy), run the backfill standalone — also
--name-keyed:
python -m rag_stack recompute-metrics --name rag_historical_surrogate
# preview without writing:
python -m rag_stack recompute-metrics --name rag_historical_surrogate --dry-run
# skip the .bak_pre_recompute backups:
python -m rag_stack recompute-metrics --name rag_historical_surrogate --no-backup
Notes:
- The default refresh is pool-only: once real evaluation trails exist they
ARE the reference; NSGA-II is only the bootstrap for a brand-new benchmark
(
--no-pool). - The pool is scoped to a single
<name>subfolder, so different function configs never cross-mix as long as they each get a distinctname. - To rebuild from scratch, delete the per-name subfolder's cache
(
rm <output_dir>/<name>/_pareto_cache.npy) and rerun with--no-pool.