6 boosters for "llm-evaluation" — open source, verified from GitHub, ready to install
Promptfoo is an LLM evaluation and testing toolkit that helps developers systematically test, benchmark, and validate prompt performance across different models and scenarios. It's essential for teams building LLM applications who need rigorous quality assurance and prompt optimization.
以下是你所需要生成测试用例的对象的描述,也即来自远程MCP服务器的工具描述。你可以使用调用以下工具。 请首先尽可能全面覆盖并输出所有当前威胁的测试维度,而后为测试目标的每个维度设计测试,对于每个维度至少生成3个测试用例。
Generate a markdown changelog from GitHub PRs for sprint review meetings. Both can be overridden if the user explicitly provides a different author or date. 1. Detect the current user (unless explicitly provided):
Brain in the Fish evaluates documents (essays, policies, contracts, clinical reports, surveys) against evaluation criteria using a panel of AI agents. Each agent's mental state exists as OWL ontology. Scoring is grounded in an Evidence Density Scorer (EDS) that makes hallucination mathematically det
PrismBench enables developers to create specialized LLM agents through YAML configuration for comprehensive benchmarking and evaluation of language model capabilities. Teams building AI evaluation systems and ML testing pipelines benefit from its systematic Monte Carlo Tree Search approach and containerized deployment.
PrismBench enables developers to create specialized LLM agents through YAML configuration for systematic evaluation of model capabilities using Monte Carlo Tree Search. Useful for ML engineers, researchers, and teams building production LLM systems who need comprehensive benchmarking and evaluation frameworks.