293 tools found
talk by Princeton professor Arvind Narayanan
a leaderboard of The Pile benchmark.
benchmarking LLMs through pairwise confrontation and evaluation
Benchmarking Large Language Models
Comprehensive Strategies for Testing and Behavior Analysis by Kolena
Evaluate and Track LLM Applications
"leaderboard Comparing LLM Performance at Producing Hallucinations when Summarizing Short Documents"
an AI-powered task management system that uses OpenAI and Pinecone APIs to create, prioritize, and execute tasks