NVIDIA · Hybrid Mamba-2 + Attention · ZeroGPU · real Triton kernels ✅
First message loads the model (~60s). ZeroGPU shared GPU — response time varies by hardware allocated.
On: model reasons before answering (slower)