40 combinatorial optimization problems wrapped in a unified
Input–Output–Tool–Heuristic (I-O-T-H) interface, evolved by
7 operators. This page is the paper's unabridged companion:
where the paper reports aggregate scores, the four tabs below expose
every problem specification, every prompt template (template + one
rendered example), all 400 final Python heuristics, and every
per-instance evaluation behind the headline numbers.
Results overview
Fig 1 is the framework schematic (Stage one I-O-T
interface → Stage two evolve / evaluate / select loop), and
Fig 2 – 5 are interactive versions
of the paper's headline result figures: aggregate ranking on 36
CO-Bench problems, zero-shot generalization to 4 novel problems with
5-seed distributions, tool-library and complementary-portfolio
ablations, and a four-backbone robustness comparison. In the
interactive figures, hover any value for the exact number; click a
column header to sort; click a chip to hide a family or component;
click a LACE row in Fig 3 to jump to that problem's per-instance
scores.
LACE framework architecture
The end-to-end pipeline behind every number on this page.
Stage one turns a natural-language CO problem
into a validated I, O, T interface and seeds an initial
portfolio H0. Stage two
evolves the portfolio with five generation operators (LR / RR / CC /
CS / DI) plus two reactive repair operators (ER / EI), evaluates
all 2n candidates, and re-selects n heuristics
by minimizing the mean rank of the best-covering heuristic per
instance, repeating until Nmax iterations.
aAggregated performance across the 36 CO-Bench problems
9 methods scored on Avg Score (primal-gap-style), Valid Rate
(fraction of problems feasible on every instance), and Survival
Rate (fraction of instances reaching ≥ 99% of the reference).
Click any column header to sort.
bAvg Score disaggregated by problem family
LACE attains the highest score in every family. Bars show the
mean Avg Score across the problems constituting each family.
Hover for exact values.
Zero-shot generalization to 4 novel problems
15 methods × 4 novel problems; each row shows the 5-seed
distribution as a box plot (dots = individual seed means, vertical
bar = mean across seeds). The 5 LLM-agent baselines marked
fail
produced no feasible algorithm on any of the four problems.
Click any LACE row to drill into per-instance scores.
aTool library reduces heuristic construction failures
Failure rate is the fraction of generated heuristics that produced
no feasible solution. Curves show the rate as a function of the
number of domain-specific tools K (0 – 15)
for the six problem families. Click a chip to hide a family.
Avg Score averaged over the test instances and problems of each
family, as a function of portfolio size n
(1 – 10).
aPer-problem performance and cost across 4 backbones
LACE driven by four frontier LLMs on the 40 problems. Best value
in each metric is highlighted. Avg Score / Valid / Survival are
higher-is-better; Tokens, Cost, and Wall-clock are
lower-is-better.
bWall-clock time decomposition
Per-problem average time broken into Generation (5 generation
operators), Repair (2 reactive repair operators), Evaluation
(heuristic execution on training instances), and Selection
(complementary portfolio selection). Evaluation dominates across
all backbones.
Problem specs
Grouped by category. Click a category to expand it, then click a problem
to see six sections in order: Description,
Input schema (input data structure),
Output schema (return spec),
Tool library (reusable helper functions),
Heuristic skeleton (the I-O-T-H-wrapped solve()
skeleton shown to the LLM), and
Heuristic example (one LLM-generated solve()).
Loading…
LLM operators
Stage One has a Heuristic Generator
that seeds the initial portfolio from scratch (per paper §4.2).
Stage Two has seven LLM operators —
five generation (LR, RR, CC, CS, DI) plus two reactive repair (ER, EI)
— that mutate, recombine, or repair existing heuristics. All
templates carry placeholders for the problem spec and the I-O-T-H
interface; Stage Two templates additionally splice in parent code.
To make the placeholders concrete, every operator is rendered
end-to-end on one example problem
(…) below, so the exact LLM-facing
text is visible.
Loading…
Benchmark scores
Per-problem evaluation outputs. The overview table summarises all 40
problems (click any column header to sort); expand a problem for its
per-instance score table. The 4 novel problems additionally
show a head-to-head comparison against multiple baselines.
Loading…
Final heuristics
The 10 LLM-generated solve() functions surviving Stage Two
evolution for each problem (400 algorithms in total). Expand a problem to
list its heuristics #1–#10; click any heuristic to view its full Python
source.