mirror of
https://github.com/discourse/discourse.git
synced 2026-08-09 21:45:25 +08:00
Implemented an eval “comparison matrix” that lets you run the same evals across multiple personas or multiple LLMs and have a judge model declare a winner with per-candidate scores. The CLI adds --compare personas|llms, keeps persona selection (auto-prepending default for persona mode), and always ensures a judge is configured. A dedicated ComparisonRunner reuses Workbench results to build candidate outputs and sends them to Judge#compare, which crafts a rubric-aware comparison prompt and parses structured winner/ratings JSON. Outputs are streamed to the console and individual run logs still get written. README documents how to use the new flag and what each mode does. |
||
|---|---|---|
| .. | ||
| ai_helper_spec.rb | ||
| discoveries_spec.rb | ||
| hyde_spec.rb | ||
| inference_spec.rb | ||
| translation_spec.rb | ||