| Kimi K3 Architecture Overview and Notes(sebastianraschka.com) | |
| 441 points by ModelForge 19 hours ago | 95 comments | |
tl;dr: Kimi K3 is a scaled-up (2.8T parameter) production version of Kimi Linear, now the largest open-weight model, incorporating efficiency-focused components like LatentMoE (down-projected MoE, similar to Nemotron 3 Ultra), multi-head latent attention, and Kimi Delta Attention. Notably, it drops RoPE entirely in favor of NoPE across all layers—a first for a frontier-level model—and adds attention residuals that weight cross-layer residual connections via attention scores for ~4% training cost. It also introduces native multimodal support. | |
HN Discussion:
| |