Skip to content
TECH CEO Daily
AIAnalysis

Self-hosted AI: Aleph Alpha's 78B Kolibri, Ai2's Olmo-core 3 and Iterate.ai's Lifeboat

An open-weight model that runs on two GPUs, an open training framework and a confidential inference engine widen the options for keeping AI workloads in-house.

By · Editor

· 4 min read · Fact-checked

The 60-second brief

  • 1Aleph Alpha's Kolibri, released October 3 under Apache 2.0, has 78.1 billion parameters and can run on two Nvidia H100s.
  • 2Ai2 says Olmo-core 3 processed 52,000 tokens per second per GPU training a 47-billion-parameter model, 2.7 times its earlier stack.
  • 3Iterate.ai says Lifeboat fits two to six times more agent sessions per GPU, based on its own tests; verify on your workloads.

The news

Self-hosted AI got three new building blocks between October 1 and October 5: Aleph Alpha's open-weight Kolibri model, the Allen Institute for AI's Olmo-core 3 training framework and Iterate.ai's Lifeboat inference engine. Together they give regulated companies more ways to keep AI workloads in-house.

Aleph Alpha, a German AI company, released Kolibri on October 3 under the Apache 2.0 license, with full weights on Hugging Face. It is a mixture-of-experts (MoE) model, a design that activates only a few specialist sub-networks for each token: 78.1 billion parameters in total, about 3.46 billion active per token, English and German support and a context window of up to 1 million tokens.

The company said it trained Kolibri on 768 Nvidia B200 GPUs on infrastructure in Germany and Finland, using about 24 trillion tokens. Its model card lists two Nvidia H100 or A100 80GB GPUs, or a single H200, B200 or B300, as the minimum to run it, and Aleph Alpha said it can serve 18 concurrent 256,000-token requests on two H100s. The company reported 96.9% on the AIME 2025 math benchmark against 84.6% for Qwen3.6-35B-A3B, while the two were close on HumanEval+ coding, 92.7% to 92.8%.

The Allen Institute for AI (Ai2), a Seattle-based research group, published Olmo-core 3 on October 1 as open code on GitHub. The framework is built to train large MoE models more efficiently. Ai2 said a 47-billion-parameter MoE processed 52,000 tokens per second per GPU on eight Nvidia B300 GPUs, about 2.7 times its earlier implementation, and that a lower-precision number format called MXFP8 lifted training throughput about 21% over the standard BF16 format.

Iterate.ai's Lifeboat, which SiliconANGLE reported on October 5 as generally available, is an inference engine, the software that serves a trained model's answers, with confidential computing built in. Model weights stay sealed and encrypted inside a hardware-based trusted execution environment, whether on a cloud confidential virtual machine or customer-owned hardware, according to the report. Iterate.ai says the engine fits two to six times as many concurrent AI agent sessions on each GPU.

In Iterate.ai's own testing on a single Nvidia RTX PRO 6000 Blackwell GPU running a Qwen 30B-A3B model, Lifeboat held 2,048 concurrent sessions, double the number with its optimizations off, and produced 8,714 tokens per second against 4,965. A free developer license covers noncommercial and evaluation use; a standard license costs $49.99 a month and the confidential computing edition $499.99 a month.

The numbers

Kolibri total parameters
78.1 billion
Kolibri parameters active per token
About 3.46 billion
Kolibri maximum context window
1 million tokens
Olmo-core 3 throughput, 47B MoE (per GPU, Ai2)
52,000 tokens/sec
Lifeboat sessions on one GPU (Iterate.ai test)
2,048
Lifeboat confidential computing edition
$499.99 a month

Why CEOs should care

For CISOs and compliance chiefs at banks, hospitals and government contractors, these releases widen the options for keeping sensitive prompts and documents off third-party AI services. Kolibri's Apache 2.0 license is permissive, and its model card says two H100s are enough to run it. Ask what GPUs you already own, and test whether a model of this size meets your accuracy bar on your own documents, not just on the vendor's benchmarks.

For CFOs, the pitch is utilization. Iterate.ai CEO Jon Nordmark argues that enterprises should find out what their existing GPUs can do before buying more for AI agents. The two-to-six-times claim comes from the company's own tests on one GPU and one model, so ask for a trial on your workloads and measure sessions per GPU and response times before changing hardware budgets.

Boards should keep the limits in view. Self-hosting moves patching, monitoring and safety controls onto your own team, and Kolibri's model card warns that it cannot replace application-level safeguards and is not recommended for unsupervised high-stakes decisions. Ai2 aims Olmo-core 3 at researchers and smaller labs priced out of large-model training; for most enterprises, the nearer decision is which open model to serve, not whether to train one.

The bigger picture

The releases fit a broader push to make AI less dependent on a handful of cloud providers. Aleph Alpha calls Kolibri sovereign: built in Germany, trained in Germany and Finland under European and German law, and designed with the EU AI Act and GDPR in mind, according to the company. Ai2 says the compute cost of training large models puts development out of reach for many academic researchers and smaller labs.

MoE designs sit at the center of that efficiency story. Kolibri uses about 3.46 billion of its 78.1 billion parameters per token, and Lifeboat loads experts selectively on MoE models, SiliconANGLE reported. Fewer active parameters per token generally mean less computing per answer, which is why open MoE models are practical candidates for in-house deployment.

What’s next

Watch for independent benchmarks of Kolibri beyond Aleph Alpha's own results, for the next-generation Olmo models Ai2 says will use Olmo-core 3, and for third-party tests of Lifeboat's session claims. Buyers weighing self-hosted AI should also check which models and GPUs each tool supports, and put confidential computing features through their own security review.

What “Fact-checked” means

Fact-checking means testing a story’s facts against the evidence before it is published. This story went through at least two separate checks before this version was published.

What we checked
Its names, figures, dates, job titles, quotes and who said what were checked against the story’s sources, including its main source where it could be opened. The headline was checked for accuracy and overstatement.
How
A first check reviewed the whole story. If it passed, a second, skeptical check went back to the sources to look for mistakes in the most important facts. If a check flagged the story, it was edited to fix the problems found, and a separate re-check then reviewed the whole story again.
Who
The checks are made by our newsroom, as steps kept separate from the writing, under rules set by our editor, . A story the checks still flag is held for the editor, who decides whether it is fixed, published or dropped.
If something is wrong
“Fact-checked” does not mean error-free. If a material error is found after publication, we correct the story and add a note saying what changed. Report an error

How we fact-check →

Companies in this story

Aleph AlphaAllen Institute for AIIterate.aiOpen-weight modelsConfidential computing

Earlier coverage of Hugging Face

All Hugging Face coverage →

Written by

Editor · Technology & Business Writer

Hussein is a writer and business technology enthusiast focused on the intersection of technology, entrepreneurship, finance, artificial intelligence, and digital innovation.

CoversAICybersecurityBig TechSaaSStartupsFintech

About this story. Researched from primary sources whenever they are available and fact-checked before publication.

Published by Tech CEO Daily, an independent publication. Masthead · Editorial standards

Follow Tech CEO Daily on Facebook for the day’s top stories in your feed.

Free newsletters

The technology briefing for people running businesses.

Daily, weekly, bi-weekly or monthly. You choose.

How often

The Daily Brief · Monday to Saturday, 7 a.m. ET

Free forever. One click to unsubscribe. We never sell your email.