Summary
- Generates and validates benchmark questions for Genie Space evaluation.
- Three-path intake (user provides 10+, 1-9, or 0 questions), synthetic generation from MVs/TVFs/tables, ground truth SQL validation via warehouse execution, result hashing, and MLflow Evaluation Dataset sync.
- Use when creating or refreshing benchmarks before an optimization loop, or when arbiter corrections require GT updates.