2,357 tools found
"leaderboard Comparing LLM Performance at Producing Hallucinations when Summarizing Short Documents"
Comprehensive Strategies for Testing and Behavior Analysis by Kolena
Benchmarking Large Language Models
benchmarking LLMs through pairwise confrontation and evaluation
an interactive eBook that compiles an extensive collection of research papers on large language model (LLM)-based multi-agent systems
"GPTeam uses GPT-4 to create multiple agents who collaborate to achieve predefined goals"
LLM Comparator is an interactive data visualization tool for evaluating and analyzing LLM responses side-by-side, developed by the PAIR team.
AI multi-agent problem solving
Educational framework exploring ergonomic, lightweight multi-agent orchestration. Managed by OpenAI Solution team.
"An Open Source Language Model Specialized in Evaluating Other Language Models."
The LLM Evaluation Framework
LLM-powered multiagent persona simulation for imagination enhancement and business insights
JARVIS, a system to connect LLMs with ML community
research on evaluation of LLMs conducted by Microsoft Research and other collaborated institutes. (Updated at: 2023/10)