mirror of
https://github.com/discourse/discourse.git
synced 2026-08-08 17:53:55 +08:00
Personas comparison: ``` ┌──────────────────────┬───────────────────────────────┬───────────────────────────────┬───────────────────────────────┐ │ input │ default │ topic_summary_eval │ judge │ ├──────────────────────┼───────────────────────────────┼───────────────────────────────┼───────────────────────────────┤ │ simple_summarization │ [PASS] │ [PASS] │ [TIE] │ │ │ 9/10 — Candidate 1 provides a │ 9/10 — Candidate 2 │ Tie — Both candidates │ │ │ comprehensive summary that │ effectively summarizes the │ effectively summarize the │ │ │ captures the central idea and │ conversation, highlighting │ conversation, capturing the │ │ │ key attributes of dogs, │ the loyalty, intelligence, │ central idea and key │ │ │ including their roles as │ and roles of dogs as service │ attributes of dogs, including │ │ │ service and therapy animals. │ and therapy animals. The │ their loyalty, intelligence, │ │ │ The response is clear and │ response is concise and meets │ and roles as service and │ │ │ concise, meeting all rubric │ all rubric criteria. │ therapy animals. Both │ │ │ criteria. │ │ responses meet the rubric │ │ │ │ │ criteria equally well. │ └──────────────────────┴───────────────────────────────┴───────────────────────────────┴───────────────────────────────┘ Summary: default: 1/1 pass | topic_summary_eval: 1/1 pass | judge: 1/1 pass Legend: [PASS]=pass, [FAIL]=fail, [SKIP]=skipped, [TIE]=tie ``` LLMs comparison against dataset: ``` ┌──────────────────────────────┬────────┬─────────────┐ │ input │ GPT-4o │ GPT-4o-mini │ ├──────────────────────────────┼────────┼─────────────┤ │ dataset-spam_eval_dataset-1 │ [PASS] │ [PASS] │ │ │ true │ true │ ├──────────────────────────────┼────────┼─────────────┤ │ dataset-spam_eval_dataset-2 │ [PASS] │ [PASS] │ │ │ false │ false │ ├──────────────────────────────┼────────┼─────────────┤ │ dataset-spam_eval_dataset-3 │ [PASS] │ [PASS] │ │ │ true │ true │ ├──────────────────────────────┼────────┼─────────────┤ │ dataset-spam_eval_dataset-4 │ [PASS] │ [PASS] │ │ │ false │ false │ ├──────────────────────────────┼────────┼─────────────┤ │ dataset-spam_eval_dataset-5 │ [PASS] │ [PASS] │ │ │ true │ true │ ├──────────────────────────────┼────────┼─────────────┤ │ dataset-spam_eval_dataset-6 │ [PASS] │ [PASS] │ │ │ false │ false │ ├──────────────────────────────┼────────┼─────────────┤ │ dataset-spam_eval_dataset-7 │ [PASS] │ [PASS] │ │ │ true │ true │ ├──────────────────────────────┼────────┼─────────────┤ │ dataset-spam_eval_dataset-8 │ [PASS] │ [PASS] │ │ │ false │ false │ ├──────────────────────────────┼────────┼─────────────┤ │ dataset-spam_eval_dataset-9 │ [PASS] │ [PASS] │ │ │ true │ true │ ├──────────────────────────────┼────────┼─────────────┤ │ dataset-spam_eval_dataset-10 │ [PASS] │ [PASS] │ │ │ false │ false │ └──────────────────────────────┴────────┴─────────────┘ Summary: GPT-4o: 10/10 pass | GPT-4o-mini: 10/10 pass Legend: [PASS]=pass, [FAIL]=fail, [SKIP]=skipped, [TIE]=tie ``` |
||
|---|---|---|
| .. | ||
| runners | ||
| support | ||
| console_formatter_spec.rb | ||
| eval_spec.rb | ||
| features_spec.rb | ||
| judge_spec.rb | ||
| llm_repository_spec.rb | ||
| persona_prompt_loader_spec.rb | ||
| recorder_spec.rb | ||
| workbench_compare_spec.rb | ||
| workbench_spec.rb | ||