Janus: Co-Evolving Programs and Evaluators for LLM Discovery with Expensive Evaluations
A novel framework named Janus tackles the difficulties of LLM-driven program discovery, particularly in scenarios where evaluations are costly, such as in scientific and engineering fields that necessitate high-fidelity simulations or physical experiments. This framework, described in an arXiv paper (2608.08189), employs LLMs to co-evolve target programs alongside executable proxy evaluators. Janus utilizes domain knowledge embedded in LLMs to create task-specific evaluator programs, which are then calibrated with actual results. To address distribution shifts, it evolves evaluators in tandem with target programs, employs a promotion-aligned objective for selection, and keeps region-conditioned portfolios with online credit updates. Given the fallibility of proxy predictions, Janus treats them as guidance to ensure robustness. This framework aims to lower evaluation costs while preserving the quality of discovery. The paper, recently submitted to arXiv, highlights its core methodology and is significant at the intersection of AI, machine learning, and automated scientific discovery.
Key facts
- Janus is a framework for LLM-driven program discovery under expensive evaluation budgets.
- It co-evolves target programs and executable proxy evaluators using LLMs.
- Janus calibrates evaluators using real outcomes and domain knowledge from LLMs.
- It mitigates distribution shift by evolving evaluators alongside target programs.
- Janus uses promotion-aligned objectives and region-conditioned portfolios with online credit updates.
- Proxy predictions are used only as guidance due to their fallibility.
- The framework addresses label scarcity and search-induced distribution shift.
- The paper is available on arXiv with identifier 2608.08189.
Entities
Institutions
- arXiv