tensorrt-llm

21.8k
davila7davila7

Optimizes LLM inference with NVIDIA TensorRT for maximum throughput and lowest latency. Use for production deployment on NVIDIA GPUs (A100/H100), when you need 10-100x faster inference than PyTorch, or for serving models with quantization (FP8/INT4), in-flight batching, and multi-GPU scaling.

Inference ServingTensorRT-LLMNVIDIA+8
191 days ago

serving-llms-vllm

21.8k
davila7davila7

Serves LLMs with high throughput using vLLM's PagedAttention and continuous batching. Use when deploying production LLM APIs, optimizing inference latency/throughput, or serving models with limited GPU memory. Supports OpenAI-compatible endpoints, quantization (GPTQ/AWQ/FP8), and tensor parallelism.

vLLMInference ServingPagedAttention+6
191 days ago

content-writing-thought-leadership

1.8k
openclawopenclaw

B2B content writing with daily workflows and batching systems across Sales/HR/Fintech/Ops Tech

191 days ago

langfuse-rate-limits

1.5k
Jeremylongshore Claude Code Plugins Plus Skills Langfuse Rate LimitsJeremylongshore Claude Code Plugins Plus Skills Langfuse Rate Limits

Implement Langfuse rate limiting, batching, and backoff patterns. Use when handling rate limit errors, optimizing trace ingestion, or managing high-volume LLM observability workloads. Trigger with phrases like "langfuse rate limit", "langfuse throttling", "langfuse 429", "langfuse batching", "langfuse high volume".

191 days ago

replit-performance-tuning

1.5k
Jeremylongshore Claude Code Plugins Plus Skills Replit Performance TuningJeremylongshore Claude Code Plugins Plus Skills Replit Performance Tuning

Optimize Replit API performance with caching, batching, and connection pooling. Use when experiencing slow API responses, implementing caching strategies, or optimizing request throughput for Replit integrations. Trigger with phrases like "replit performance", "optimize replit", "replit latency", "replit caching", "replit slow", "replit batch".

191 days ago

vercel-performance-tuning

1.5k
jeremylongshorejeremylongshore

Optimize Vercel API performance with caching, batching, and connection pooling. Use when experiencing slow API responses, implementing caching strategies, or optimizing request throughput for Vercel integrations. Trigger with phrases like "vercel performance", "optimize vercel", "vercel latency", "vercel caching", "vercel slow", "vercel batch".

191 days ago

firecrawl-performance-tuning

1.5k
Jeremylongshore Claude Code Plugins Plus Skills Firecrawl Performance TuningJeremylongshore Claude Code Plugins Plus Skills Firecrawl Performance Tuning

Optimize FireCrawl API performance with caching, batching, and connection pooling. Use when experiencing slow API responses, implementing caching strategies, or optimizing request throughput for FireCrawl integrations. Trigger with phrases like "firecrawl performance", "optimize firecrawl", "firecrawl latency", "firecrawl caching", "firecrawl slow", "firecrawl batch".

191 days ago

perplexity-performance-tuning

1.5k
Jeremylongshore Claude Code Plugins Plus Skills Perplexity Performance TuningJeremylongshore Claude Code Plugins Plus Skills Perplexity Performance Tuning

Optimize Perplexity API performance with caching, batching, and connection pooling. Use when experiencing slow API responses, implementing caching strategies, or optimizing request throughput for Perplexity integrations. Trigger with phrases like "perplexity performance", "optimize perplexity", "perplexity latency", "perplexity caching", "perplexity slow", "perplexity batch".

191 days ago

supabase-performance-tuning

1.5k
Jeremylongshore Claude Code Plugins Plus Skills Supabase Performance TuningJeremylongshore Claude Code Plugins Plus Skills Supabase Performance Tuning

Optimize Supabase API performance with caching, batching, and connection pooling. Use when experiencing slow API responses, implementing caching strategies, or optimizing request throughput for Supabase integrations. Trigger with phrases like "supabase performance", "optimize supabase", "supabase latency", "supabase caching", "supabase slow", "supabase batch".

191 days ago

instantly-performance-tuning

1.5k
jeremylongshorejeremylongshore

Optimize Instantly API performance with caching, batching, and connection pooling. Use when experiencing slow API responses, implementing caching strategies, or optimizing request throughput for Instantly integrations. Trigger with phrases like "instantly performance", "optimize instantly", "instantly latency", "instantly caching", "instantly slow", "instantly batch".

191 days ago

mistral-performance-tuning

1.5k
Jeremylongshore Claude Code Plugins Plus Skills Mistral Performance TuningJeremylongshore Claude Code Plugins Plus Skills Mistral Performance Tuning

Optimize Mistral AI performance with caching, batching, and latency reduction. Use when experiencing slow API responses, implementing caching strategies, or optimizing request throughput for Mistral AI integrations. Trigger with phrases like "mistral performance", "optimize mistral", "mistral latency", "mistral caching", "mistral slow", "mistral batch".

191 days ago

clay-performance-tuning

1.5k
Jeremylongshore Claude Code Plugins Plus Skills Clay Performance TuningJeremylongshore Claude Code Plugins Plus Skills Clay Performance Tuning

Optimize Clay API performance with caching, batching, and connection pooling. Use when experiencing slow API responses, implementing caching strategies, or optimizing request throughput for Clay integrations. Trigger with phrases like "clay performance", "optimize clay", "clay latency", "clay caching", "clay slow", "clay batch".

191 days ago

groq-performance-tuning

1.5k
Jeremylongshore Claude Code Plugins Plus Skills Groq Performance TuningJeremylongshore Claude Code Plugins Plus Skills Groq Performance Tuning

Optimize Groq API performance with caching, batching, and connection pooling. Use when experiencing slow API responses, implementing caching strategies, or optimizing request throughput for Groq integrations. Trigger with phrases like "groq performance", "optimize groq", "groq latency", "groq caching", "groq slow", "groq batch".

191 days ago

vastai-performance-tuning

1.5k
jeremylongshorejeremylongshore

Optimize Vast.ai API performance with caching, batching, and connection pooling. Use when experiencing slow API responses, implementing caching strategies, or optimizing request throughput for Vast.ai integrations. Trigger with phrases like "vastai performance", "optimize vastai", "vastai latency", "vastai caching", "vastai slow", "vastai batch".

191 days ago

retellai-performance-tuning

1.5k
jeremylongshorejeremylongshore

Optimize Retell AI API performance with caching, batching, and connection pooling. Use when experiencing slow API responses, implementing caching strategies, or optimizing request throughput for Retell AI integrations. Trigger with phrases like "retellai performance", "optimize retellai", "retellai latency", "retellai caching", "retellai slow", "retellai batch".

191 days ago

fireflies-performance-tuning

1.5k
Jeremylongshore Claude Code Plugins Plus Skills Fireflies Performance TuningJeremylongshore Claude Code Plugins Plus Skills Fireflies Performance Tuning

Optimize Fireflies.ai API performance with caching, batching, and connection pooling. Use when experiencing slow API responses, implementing caching strategies, or optimizing request throughput for Fireflies.ai integrations. Trigger with phrases like "fireflies performance", "optimize fireflies", "fireflies latency", "fireflies caching", "fireflies slow", "fireflies batch".

191 days ago

posthog-performance-tuning

1.5k
Jeremylongshore Claude Code Plugins Plus Skills Posthog Performance TuningJeremylongshore Claude Code Plugins Plus Skills Posthog Performance Tuning

Optimize PostHog API performance with caching, batching, and connection pooling. Use when experiencing slow API responses, implementing caching strategies, or optimizing request throughput for PostHog integrations. Trigger with phrases like "posthog performance", "optimize posthog", "posthog latency", "posthog caching", "posthog slow", "posthog batch".

191 days ago

windsurf-performance-tuning

1.5k
Jeremylongshore Claude Code Plugins Plus Skills Windsurf Performance TuningJeremylongshore Claude Code Plugins Plus Skills Windsurf Performance Tuning

Optimize Windsurf API performance with caching, batching, and connection pooling. Use when experiencing slow API responses, implementing caching strategies, or optimizing request throughput for Windsurf integrations. Trigger with phrases like "windsurf performance", "optimize windsurf", "windsurf latency", "windsurf caching", "windsurf slow", "windsurf batch".

191 days ago

ideogram-performance-tuning

1.5k
Jeremylongshore Claude Code Plugins Plus Skills Ideogram Performance TuningJeremylongshore Claude Code Plugins Plus Skills Ideogram Performance Tuning

Optimize Ideogram API performance with caching, batching, and connection pooling. Use when experiencing slow API responses, implementing caching strategies, or optimizing request throughput for Ideogram integrations. Trigger with phrases like "ideogram performance", "optimize ideogram", "ideogram latency", "ideogram caching", "ideogram slow", "ideogram batch".

191 days ago

processing-api-batches

1.5k
jeremylongshorejeremylongshore

Optimize bulk API requests with batching, throttling, and parallel execution. Use when processing bulk API operations efficiently. Trigger with phrases like "process bulk requests", "batch API calls", or "handle batch operations".

191 days ago

documenso-performance-tuning

1.5k
jeremylongshorejeremylongshore

Optimize Documenso integration performance with caching, batching, and efficient patterns. Use when improving response times, reducing API calls, or optimizing bulk document operations. Trigger with phrases like "documenso performance", "optimize documenso", "documenso caching", "documenso batch operations".

191 days ago

analyzing-network-latency

1.5k
Jeremylongshore Claude Code Plugins Plus Skills Network Latency AnalyzerJeremylongshore Claude Code Plugins Plus Skills Network Latency Analyzer

This skill enables Claude to analyze network latency and optimize request patterns within an application. It helps identify bottlenecks and suggest improvements for faster and more efficient network communication. Use this skill when the user asks to "analyze network latency", "optimize request patterns", or when facing performance issues related to network requests. It focuses on identifying serial requests that can be parallelized, opportunities for request batching, connection pooling improvements, timeout configuration adjustments, and DNS resolution enhancements. The skill provides concrete suggestions for reducing latency and improving overall network performance.

191 days ago

coderabbit-performance-tuning

1.5k
Jeremylongshore Claude Code Plugins Plus Skills Coderabbit Performance TuningJeremylongshore Claude Code Plugins Plus Skills Coderabbit Performance Tuning

Optimize CodeRabbit API performance with caching, batching, and connection pooling. Use when experiencing slow API responses, implementing caching strategies, or optimizing request throughput for CodeRabbit integrations. Trigger with phrases like "coderabbit performance", "optimize coderabbit", "coderabbit latency", "coderabbit caching", "coderabbit slow", "coderabbit batch".

191 days ago

exa-performance-tuning

1.5k
Jeremylongshore Claude Code Plugins Plus Skills Exa Performance TuningJeremylongshore Claude Code Plugins Plus Skills Exa Performance Tuning

Optimize Exa API performance with caching, batching, and connection pooling. Use when experiencing slow API responses, implementing caching strategies, or optimizing request throughput for Exa integrations. Trigger with phrases like "exa performance", "optimize exa", "exa latency", "exa caching", "exa slow", "exa batch".

191 days ago

bio-batch-downloads

293
GPTomicsGPTomics

Download large datasets from NCBI efficiently using history server, batching, and rate limiting. Use when performing bulk sequence downloads, handling large query results, or production-scale data retrieval.

191 days ago

batching-patterns

241
MadAppGangMadAppGang

Batch all related operations into single messages for maximum parallelism and performance. Use when launching multiple agents, reading multiple files, running parallel searches, optimizing workflow speed, or avoiding sequential execution bottlenecks. Trigger keywords - "batching", "parallel", "single message", "golden rule", "concurrent", "performance", "sequential bottleneck", "speed optimization".

orchestrationbatchingparallel+3
191 days ago

quiz-generator

150
panaversitypanaversity

Generate 50-question interactive quizzes using the Quiz component with randomized batching. Use when creating end-of-chapter assessments. Displays 15-20 questions per session with immediate feedback. NOT for static markdown quizzes.

191 days ago

absinthe-resolvers

98
TheBushidoCollectiveTheBushidoCollective

Use when implementing GraphQL resolvers with Absinthe. Covers resolver patterns, dataloader integration, batching, and error handling.

191 days ago

graphql-performance

98
TheBushidoCollectiveTheBushidoCollective

Use when optimizing GraphQL API performance with query complexity analysis, batching, caching strategies, depth limiting, monitoring, and database optimization.

191 days ago

graphql-resolvers

98
TheBushidoCollectiveTheBushidoCollective

Use when implementing GraphQL resolvers with resolver functions, context management, DataLoader batching, error handling, authentication, and testing strategies.

191 days ago

Unity Scene Optimizer

75
Dev-GOMDev-GOM

Analyzes scenes for performance bottlenecks (draw calls, batching, textures, GameObjects). Use when optimizing scenes or investigating performance issues.

191 days ago

starkzap-sdk

73
keep-starknet-strangekeep-starknet-strange

Use when integrating or maintaining applications built with keep-starknet-strange/starkzap. Covers StarkSDK setup, onboarding (Signer/Privy/Cartridge), wallet lifecycle, sponsored transactions, ERC20 transfers, staking flows, tx builder batching, examples, tests, and generated presets.

191 days ago

parallel-execution-optimizer

70
marcusgollmarcusgoll

Identify and execute independent operations in parallel for 3-5x speedup. Auto-analyzes task dependencies, groups into batches, launches parallel Task() calls. Applies to /optimize (5 checks), /ship pre-flight (5 checks), /implement (task batching), /prototype (N screens). Auto-triggers when detecting multiple independent operations in a phase.

191 days ago

llm-inference-batching-scheduler

62
letta-ailetta-ai

Guidance for implementing batching schedulers for LLM inference systems with compilation-based accelerators. This skill applies when optimizing request batching to minimize cost while meeting latency thresholds, particularly when dealing with shape compilation costs, padding overhead, and multi-bucket request distributions. Use this skill for tasks involving batch planning, shape selection, generation-length bucketing, and cost-model-driven optimization for neural network inference.

191 days ago

remote-functions

58
spences10spences10

Use when building, auditing, or reviewing SvelteKit remote functions for validation, batching, and optimistic UI patterns

191 days ago

ai-llm-inference

34
vasilyu1983vasilyu1983

Operational patterns for LLM inference: latency budgeting, tail-latency control, caching, batching/scheduling, quantization/compression, parallelism, and reliable serving at scale. Emphasizes production-grade performance, cost control, and observability.

191 days ago

focus-timeboxing-8020

31
lyndonkllyndonkl

Use when managing time and attention, combating procrastination or context-switching, prioritizing high-impact work, planning daily/weekly schedules, improving focus and productivity, or when user mentions timeboxing, Pomodoro, deep work, 80/20 rule, Pareto principle, focus blocks, task batching, energy management, or needs structured approach to getting important work done.

191 days ago

speed-of-light

30
SimHackerSimHacker

Many turns in one call. Instant communication. No round-trips.

moollmoptimizationlatency+2
191 days ago

social-inbox

14
0xAxiom0xAxiom

Aggregate, score, and prioritize social mentions across X/Twitter. Outputs a ranked inbox with engagement scores and draft context for efficient reply batching.

191 days ago

llm-cost-optimization

10
BagelHoleBagelHole

Reduce LLM API and infrastructure costs through model selection, prompt caching, batching, caching, quantization, and self-hosting strategies. Track spend by team and model, set budgets, and implement cost-aware routing.

191 days ago

vllm-server

10
BagelHoleBagelHole

Deploy and manage vLLM for high-throughput LLM inference. Configure continuous batching, tensor parallelism, quantization, and OpenAI-compatible API endpoints for production LLM serving.

191 days ago

vllm

9
TerminalSkillsTerminalSkills

High-throughput LLM serving engine with PagedAttention for efficient memory management. Serves open-source models with OpenAI-compatible API, continuous batching, tensor parallelism, and quantization support. Optimized for production inference workloads.

191 days ago

realtime-analytics

9
TerminalSkillsTerminalSkills

Build real-time analytics pipelines from scratch. Use when someone asks to "set up analytics", "build a dashboard", "track events in real time", "ClickHouse analytics", "event ingestion pipeline", or "live metrics". Covers event schema design, ingestion services with batching, ClickHouse table optimization, aggregation queries, and dashboard wiring.

191 days ago

triton

9
TerminalSkillsTerminalSkills

NVIDIA Triton Inference Server for deploying AI models at scale. Supports multiple frameworks (ONNX, TensorRT, PyTorch, TensorFlow), model ensembles, dynamic batching, model versioning, and GPU/CPU inference with high throughput and low latency.

191 days ago

processing-api-batches

8
BbgnsurfTechBbgnsurfTech

Process bulk API requests efficiently with batching, throttling, and parallel execution. Use when processing bulk API operations efficiently. Trigger with phrases like "process bulk requests", "batch API calls", or "handle batch operations".

191 days ago

llamacpp

7
datathingsdatathings

Complete llama.cpp C/C++ API reference covering model loading, inference, text generation, embeddings, chat, tokenization, sampling, batching, KV cache, LoRA adapters, and state management. Triggers on: llama.cpp questions, LLM inference code, GGUF models, local AI/ML inference, C/C++ LLM integration, "how do I use llama.cpp", API function lookups, implementation questions, troubleshooting llama.cpp issues, and any llama-cpp or ggerganov/llama.cpp mentions.

191 days ago

unity-performance

6
creator-hiancreator-hian

Optimize Unity game performance through profiling, draw call reduction, and resource management. Masters batching, LOD, occlusion culling, and mobile optimization. Use for performance bottlenecks, frame rate issues, or optimization strategies.

191 days ago

accelint-ts-performance

6
gohypergiantgohypergiant

Systematic JavaScript/TypeScript performance audit and optimization using V8 profiling and runtime patterns. Use when (1) Users say 'optimize performance', 'audit performance', 'this is slow', 'reduce allocations', 'improve speed', 'check performance', (2) Analyzing code for performance anti-patterns (O(n²) complexity, excessive allocations, I/O blocking, template literal waste), (3) Optimizing functions regardless of current usage context - utilities, formatters, parsers are often called in hot paths even when they appear simple, (4) Fixing V8 deoptimization (monomorphic/polymorphic issues, inline caching). Audits ALL code for anti-patterns and reports findings with expected gains. Covers loops, caching, batching, memory locality, algorithmic complexity fixes with ❌/✅ patterns.

191 days ago