Groq
Ultra-fast inference on LPU hardware — near-zero latency for realtime apps.
Why it's listed
If you care about ultra-low latency, Groq’s LPU hardware makes real-time inference snappy—great for chatbots, voice, and debugging. It skips the warm-up and cold starts of typical GPU services, letting you iterate faster.
In the same category
OpenRouter
One API key, hundreds of models — compare, fall back and rate-limit without the plumbing.
Ollama
The lowest bar for local open models — one command for Llama, Qwen or DeepSeek; data stays on your machine.
Hugging Face
The GitHub of models and datasets — find models, check leaderboards, run demos.
Firecrawl
Clean web-to-markdown/structured-data for LLMs — essential plumbing for content pipelines and RAG.
Found a great tool?
Submit it — after human review it joins the AI Directory.
Submit a tool