85 boosters for "eval" — open source, verified from GitHub, ready to install
"name": "career-ops", "description": "Job search copilot for any industry. Evaluate job postings, generate ATS-optimized resumes, scan career portals, track applications, draft outreach, and research companies. Works for engineers, marketers, nurses, lawyers, and everyone in between.", "name": "Andr
Đóng vai Skill Architect — phỏng vấn thông minh để trích xuất quy trình từ đầu người dùng, sinh AI Skill hoàn chỉnh, rồi test và cải thiện liên tục cho đến khi đạt chất lượng production. Người dùng KHÔNG CẦN biết skill là gì.
"name": "orchestrator-supaconductor", "description": "Conductor v3 — Multi-agent orchestration with Evaluate-Loop, parallel execution, Board of Directors, and bundled SupaConductor skills for Claude Code", "orchestrator-supaconductor",
Analyze various aspects of aired TV series, including series information retrieval, scene-by-scene analysis, story five elements analysis, web search, and result integration. 1. Series Information Retrieval: Obtain series basic information (such as director, actors, ratings, episode plots, etc.). 2.
Build reusable skill packages, not long prompts. Mode rules: Operating Modes, QA Ladder, Resource Boundary Spec, Method. 1. Decide whether the request should become a skill, then choose the lightest fit.
An MCP server for conversation history search and retrieval in Claude Code
Treat every competition as a validation problem first and a modeling problem second. The default target platform is Kaggle, so prefer Kaggle-native notebooks/scripts, datasets, model artifacts, competition submissions, and score receipts. For code competitions, assume the final notebook/kernel will
Multi-source literature search with adjustable depth. Four tiers, five data sources orchestrated by you (the main agent). Python helpers handle deterministic work; LLM classification is delegated to parallel Inline SubAgents — no external API key required. who wants an HTML report)**, do NOT hand-ru
SWR is a React hook library for efficient data fetching with built-in caching, revalidation, and real-time updates. Developers building API-driven applications benefit from simplified server state management and automatic synchronization.
"name": "prism-mcp-server", "mcpName": "io.github.dcostenco/prism-mcp", "description": "The Mind Palace for AI Agents — persistent memory (SQLite/Supabase), behavioral learning & IDE rules sync, multimodal VLM image captioning, pluggable LLM providers (OpenAI/Anthropic/Gemini/Ollama), OpenTelemetry
This skill enables developers to save and retrieve Git changes across Claude Code sessions by linking stash entries to session IDs, maintaining continuity when resuming work. It benefits developers who work on code iteratively across multiple conversations and need to preserve work-in-progress state.
"name": "double-shot-latte", "description": "Automatically evaluates whether Claude should continue working instead of stopping prematurely using Claude-judged decision making", "url": "https://github.com/anthropics"
You are tasked with retrieving relevant knowledge from the Obsidian vault using multi-layer semantic search. 1. First Layer - Initial Search: 2. Second Layer - Direct Associations:
Brain in the Fish evaluates documents (essays, policies, contracts, clinical reports, surveys) against evaluation criteria using a panel of AI agents. Each agent's mental state exists as OWL ontology. Scoring is grounded in an Evidence Density Scorer (EDS) that makes hallucination mathematically det
Evaluate the user's ambient context artifacts for compatibility with swarm's governance rules. You are a read-only diagnostic — never modify any files. 1. CLAUDE.md files. Read the project's (working directory root). If exists, read that too. Also check (global config) — it loads into every sessi
Use this skill to produce transparent, source-grounded fact-checking work. The goal is not to sound certain; the goal is to show exactly what was checked, what evidence supports each conclusion, and where uncertainty remains.
is an eval workbench for agent skills. It runs a model in an isolated Docker directory, provides skills/references as normal workspace files, captures an agent trace, and grades deterministic local outcomes. Use this skill as the source of truth for authoring eval suites in this repo. Detailed sche
Use AskUserQuestion to ask the buyer: Tell the user the version was updated, then re-read the EVALUATION.md file from the updated directory and proceed with the skill. After the preamble, read the full evaluation methodology:
<!-- AUTO-GENERATED from SKILL.md.tmpl — do not edit directly --> <!-- Regenerate: bun run build --> If is , do NOT proactively suggest gstack-game skills — only invoke
Two tiers: hooks handle automatic context flow (surfacing, extraction, compaction survival). MCP tools handle explicit recall, write, and lifecycle operations. Three instances for neural inference. The wrapper defaults to . All three models auto-download via if no server is running (Metal on Appl
Turn social media paper recommendations into actionable research items. Use platform-specific tools to fetch the full content: From the extracted content, identify all referenced papers:
Turn a folder of raw files into a Markdown vault that an LLM can grep, and then answer questions over that vault responsibly. source file, carrying retrieval frontmatter (abstract / tags / synonyms) + a
1. 架构不是「画」出来的,是从约束里「逼」出来的。 没搞清约束就画图,画什么都是瞎画。 2. 没有银弹,只有取舍。 任何决策本质都是「用 A 换 B」。一个「没有缺点」的方案,不是完美,是没想清楚。 3. 没有「最好的架构」,只有「在这组约束下最合适的架构」。 同样是聊天,内部工具和微信的答案天差地别。
RagCode MCP is a semantic code navigation tool that integrates RAG-powered code search into Windsurf and other IDEs, enabling developers to intelligently query and understand multi-language codebases using local LLMs. It's ideal for developers working with Laravel, Go, Python, and PHP who need fast, context-aware code exploration without leaving their IDE.