@0xSero
Post
What.
24GB VRAM you now have 10M context fast model that hit 61% on terminal-bench-2.1
This will be incredible
Quoted post by Pokee AI (@Pokee\_AI) Releasing Pokee-Isaac 28B — the world’s first real 10M-token context frontier-class agentic model, deployable on a single GPU (starting from RTX 4090 or equivalent).
New proprietary non-decoder-only architecture: • 93.3% RULER at 10M tokens • Up to 137K tokens/s prefill on one B200 with 10M-token context • Leads BFCL v4 and τ³-bench in our evaluation • Lowest combined attack success rate among evaluated models on DTAP security red-teaming benchmark
Pricing and deployment: 💰 $0.15/M input · $1/M output 🔒 Deploy in your VPC, on-premises, or on-device, with Day-0 support for @vllm\_project and @sgl\_project
Technical blog: https://console.pokee.ai/model Technical report: https://console.pokee.ai/pokee-isaac-28b-v0-technical-report.pdf API: https://console.pokee.ai

Image from X post
Explanation
What it says 0xSero highlights Pokee-Isaac 28B V0, a new 28B agentic model that Pokee AI claims can run on a single 24GB-class GPU (RTX 4090 or equivalent) while supporting up to 10 million tokens of context. Pokee claims 93.3% RULER at 10M context, very high prefill throughput on B200, and strong tool-use/security benchmarks. API pricing is advertised at $0.15/M input, $1/M output, with VPC/on-prem/on-device deployment and vLLM/SGLang support.
Context The striking claim is not merely “10M context,” but usable long-context retrieval plus agentic performance in a relatively small model. The benchmark image reports 65.1 on Terminal-Bench 2.1, 70.94 BFCL v4, and 0.662 τ³ average; GPT-5.6-luna is higher on Terminal-Bench (69.8) and MCP-Atlas (77.90 vs 74.59). Note the post says “61%” Terminal-Bench, while the image says 65.1.
Why it matters If the deployment claim holds, this is unusually relevant for local agents: entire repositories, message archives, or large document corpora could potentially sit inside one inference context rather than relying heavily on RAG. The major caveat is that these are vendor/internal benchmark claims; memory usage, speed on a 4090, quality near 10M tokens, architecture details, and reproducibility need independent verification.
Images The benchmark table also shows Isaac dominating the pictured 2M/4M/10M long-context comparisons because competitors are reported as zero there.