17 boosters for "inference" — open source, verified from GitHub, ready to install
Build and publish a Gradio demo on Hugging Face Spaces that runs inference with a user-provided LoRA. Use whenever someone asks to create, generate, ship, or publish "a Space", "a demo", "a Gradio app", or "a playground" for a LoRA — whether the base model is Qwen-Image, Qwen-Image-Edit, LTX, or ano
Provides the Hugging Face Hub CLI (`hf`) tool for downloading, uploading, and managing models, datasets, and Spaces directly from Claude Code. Essential for developers integrating Hugging Face resources into AI workflows.
Hugging Face Spaces host machine-learning applications. There are 1M+ today; each Space is a git repo. This skill covers creating, building, debugging, and maintaining them. Before anything else: 1. Check the CLI is installed: . If not, .
estimates the required memory for inference, including model weights and an optional KV cache, for Safetensors and GGUF for models on the Hugging Face Hub using HTTP Range requests i.e., without downloading or loading any weights locally. Run with pointing to the Hugging Face Hub repository which w
Run any workload on fully managed Hugging Face infrastructure. No local setup required—jobs run on cloud CPUs, GPUs, or TPUs and can persist results to the Hugging Face Hub. Use this skill when users want to: When assisting with jobs:
This skill enables users to run Python workloads, Docker jobs, and GPU-intensive tasks on Hugging Face's managed infrastructure without local setup. It's valuable for ML engineers, data scientists, and developers needing cloud compute for training, inference, and batch processing.
A skill for adapting and optimizing Hugging Face or custom LLM models to run efficiently on vLLM with Ascend NPU support, enabling developers to validate and deploy models with deterministic testing and single-commit delivery.
A hands-on course teaching developers how to build AI-powered agents using Foundry Local with modern SDK patterns, function calling, and hybrid cloud designs. Benefits beginners and intermediate developers wanting to prototype agentic applications quickly on edge devices.
1. Run InterProScan for domain/family annotation. 2. Run eggnog-mapper for orthology-based annotation. 3. Run DIAMOND and resolve taxonomy with TaxonKit.
Graphsignal observes inference workloads from a sidecar process — the profiler. It never shares a process with CUDA: the profiler watches the workload externally via CUPTI, OTLP/gRPC, Prometheus scraping, and NVML. Auto-instrumentation covers vLLM, SGLang, and PyTorch out of the box. Two install pat
"description": "Modern R development skills for Claude Code - tidyverse patterns, rlang metaprogramming, Bayesian inference, performance optimization, and more", "name": "Claude Code R Skills" "homepage": "https://github.com/ab604/claude-code-r-skills",
A practical guide for deploying serverless Python applications on Modal, enabling developers to run GPU-accelerated AI/ML workloads, web APIs, and batch jobs with minimal infrastructure configuration.
"name": "inference-builder", "description": "Generate deployable Vision AI pipelines with high-performance video and streaming capabilities on NVIDIA GPUs. Create GPU-accelerated inference microservices using DeepStream, Triton, vLLM, TensorRT-LLM, or PyTorch backends.",
"$schema": "https://anthropic.com/claude-code/marketplace.schema.json", "name": "XiaoConstantine", "email": "constantine124@gmail.com"
"description": "Save tokens and cut inference costs by routing compute through MVM nodes on the cheapest available energy.", "url": "https://github.com/nhevers" "repository": "https://github.com/nhevers/mica-plugin",
"name": "@gpu-bridge/mcp-server", "description": "GPU-Bridge MCP Server — 30 AI services as MCP tools. LLM, image, video, audio, embeddings, reranking, PDF parsing, NSFW detection & more. x402 native for autonomous agents.", "gpu-bridge-mcp": "index.js"
"id": "ac.inference.sh/mcp", "name": "inference.sh", "description": "Run 150+ AI apps — image, video, audio, LLMs, 3D and more. Browse, execute, stream results."