| Mercury 2.5(inceptionlabs.ai) | |
| 245 points by Topfi 1 day ago | 52 comments | |
tl;dr: Inception has released Mercury 2.5, claimed to be the largest diffusion-based LLM ever trained, delivering 1,107 tokens/sec on NVIDIA GPUs with a 260K context window and pricing at $0.20/$0.75 per million input/output tokens (80% off at launch). The model reportedly matches cost-optimized frontier models like GPT-5.6 Luna, Gemini 3.5 Flash-Lite, and Claude Haiku 4.5, with production users citing sub-200ms latencies for voice agents and 82% latency reductions for coding context compaction. Inception also previewed Mercury Voice (sub-170ms TTFT) and Mercury Router for model routing. | |
HN Discussion:
| |