low latency AI Agent Skills
Browse 2 skills related to low latency
tensorrt-llm
21.8k
Optimizes LLM inference with NVIDIA TensorRT for maximum throughput and lowest latency. Use for production deployment on NVIDIA GPUs (A100/H100), when you need 10-100x faster inference than PyTorch, or for serving models with quantization (FP8/INT4), in-flight batching, and multi-GPU scaling.
Inference ServingTensorRT-LLMNVIDIA+8
142 days ago
pinecone
21.8k
Managed vector database for production AI applications. Fully managed, auto-scaling, with hybrid search (dense + sparse), metadata filtering, and namespaces. Low latency (<100ms p95). Use for production RAG, recommendation systems, or semantic search at scale. Best for serverless, managed infrastructure.
RAGPineconeVector Database+7
142 days ago