The news
Self-hosted AI got three new building blocks between October 1 and October 5: Aleph Alpha's open-weight Kolibri model, the Allen Institute for AI's Olmo-core 3 training framework and Iterate.ai's Lifeboat inference engine. Together they give regulated companies more ways to keep AI workloads in-house.
Aleph Alpha, a German AI company, released Kolibri on October 3 under the Apache 2.0 license, with full weights on Hugging Face. It is a mixture-of-experts (MoE) model, a design that activates only a few specialist sub-networks for each token: 78.1 billion parameters in total, about 3.46 billion active per token, English and German support and a context window of up to 1 million tokens.
The company said it trained Kolibri on 768 Nvidia B200 GPUs on infrastructure in Germany and Finland, using about 24 trillion tokens. Its model card lists two Nvidia H100 or A100 80GB GPUs, or a single H200, B200 or B300, as the minimum to run it, and Aleph Alpha said it can serve 18 concurrent 256,000-token requests on two H100s. The company reported 96.9% on the AIME 2025 math benchmark against 84.6% for Qwen3.6-35B-A3B, while the two were close on HumanEval+ coding, 92.7% to 92.8%.
The Allen Institute for AI (Ai2), a Seattle-based research group, published Olmo-core 3 on October 1 as open code on GitHub. The framework is built to train large MoE models more efficiently. Ai2 said a 47-billion-parameter MoE processed 52,000 tokens per second per GPU on eight Nvidia B300 GPUs, about 2.7 times its earlier implementation, and that a lower-precision number format called MXFP8 lifted training throughput about 21% over the standard BF16 format.
Iterate.ai's Lifeboat, which SiliconANGLE reported on October 5 as generally available, is an inference engine, the software that serves a trained model's answers, with confidential computing built in. Model weights stay sealed and encrypted inside a hardware-based trusted execution environment, whether on a cloud confidential virtual machine or customer-owned hardware, according to the report. Iterate.ai says the engine fits two to six times as many concurrent AI agent sessions on each GPU.
In Iterate.ai's own testing on a single Nvidia RTX PRO 6000 Blackwell GPU running a Qwen 30B-A3B model, Lifeboat held 2,048 concurrent sessions, double the number with its optimizations off, and produced 8,714 tokens per second against 4,965. A free developer license covers noncommercial and evaluation use; a standard license costs $49.99 a month and the confidential computing edition $499.99 a month.
The numbers
- Kolibri total parameters
- 78.1 billion
- Kolibri parameters active per token
- About 3.46 billion
- Kolibri maximum context window
- 1 million tokens
- Olmo-core 3 throughput, 47B MoE (per GPU, Ai2)
- 52,000 tokens/sec
- Lifeboat sessions on one GPU (Iterate.ai test)
- 2,048
- Lifeboat confidential computing edition
- $499.99 a month
Why CEOs should care
For CISOs and compliance chiefs at banks, hospitals and government contractors, these releases widen the options for keeping sensitive prompts and documents off third-party AI services. Kolibri's Apache 2.0 license is permissive, and its model card says two H100s are enough to run it. Ask what GPUs you already own, and test whether a model of this size meets your accuracy bar on your own documents, not just on the vendor's benchmarks.
For CFOs, the pitch is utilization. Iterate.ai CEO Jon Nordmark argues that enterprises should find out what their existing GPUs can do before buying more for AI agents. The two-to-six-times claim comes from the company's own tests on one GPU and one model, so ask for a trial on your workloads and measure sessions per GPU and response times before changing hardware budgets.
Boards should keep the limits in view. Self-hosting moves patching, monitoring and safety controls onto your own team, and Kolibri's model card warns that it cannot replace application-level safeguards and is not recommended for unsupervised high-stakes decisions. Ai2 aims Olmo-core 3 at researchers and smaller labs priced out of large-model training; for most enterprises, the nearer decision is which open model to serve, not whether to train one.
The bigger picture
The releases fit a broader push to make AI less dependent on a handful of cloud providers. Aleph Alpha calls Kolibri sovereign: built in Germany, trained in Germany and Finland under European and German law, and designed with the EU AI Act and GDPR in mind, according to the company. Ai2 says the compute cost of training large models puts development out of reach for many academic researchers and smaller labs.
MoE designs sit at the center of that efficiency story. Kolibri uses about 3.46 billion of its 78.1 billion parameters per token, and Lifeboat loads experts selectively on MoE models, SiliconANGLE reported. Fewer active parameters per token generally mean less computing per answer, which is why open MoE models are practical candidates for in-house deployment.
What’s next
Watch for independent benchmarks of Kolibri beyond Aleph Alpha's own results, for the next-generation Olmo models Ai2 says will use Olmo-core 3, and for third-party tests of Lifeboat's session claims. Buyers weighing self-hosted AI should also check which models and GPUs each tool supports, and put confidential computing features through their own security review.
What “Fact-checked” means
Fact-checking means testing a story’s facts against the evidence before it is published. This story went through at least two separate checks before this version was published.
- What we checked
- Its names, figures, dates, job titles, quotes and who said what were checked against the story’s sources, including its main source where it could be opened. The headline was checked for accuracy and overstatement.
- How
- A first check reviewed the whole story. If it passed, a second, skeptical check went back to the sources to look for mistakes in the most important facts. If a check flagged the story, it was edited to fix the problems found, and a separate re-check then reviewed the whole story again.
- Who
- The checks are made by our newsroom, as steps kept separate from the writing, under rules set by our editor, Hussein Mukhtar. A story the checks still flag is held for the editor, who decides whether it is fixed, published or dropped.
- If something is wrong
- “Fact-checked” does not mean error-free. If a material error is found after publication, we correct the story and add a note saying what changed. Report an error
Companies in this story








