HARNESS AA SERVICE

AI that gets better every week it runs

Brainsless is the endpoint your agent ships behind, with staging and CI/CD for AI behavior built in.

Built by the team behind AI serving 2.1 million students. Now running our own product, Planless.

How it works

It drills your agent through the World Compiler, turns every failure into a test, generates the fix, and ships it live as a tested version with one-click rollback. The weights never move. Only the harness does.

Connect

Connect your repo and answer a few short questions. We analyze it: how your agent runs, the tools it uses. That's what feeds the World Compiler. The whole platform also takes plain English, just describe it instead.

World Compiler

Your repo compiles into a world: working copies of your systems, same endpoints, same data, users who act like yours. It's the environment your agent runs and drills in, and it recompiles whenever your code or traffic changes.

Drill

Your agent runs that world thousands of times a night, scored on what it decides, claims, and reaches for. Grade one and it becomes a standard every version must pass. Drills pile up where it broke last, until it holds.

Ship

One API behind your product. No SDK, no prompt files, no eval scripts. Versions, diffs, and rollback come built in.

CI/CD FOR AI BEHAVIOR

Ship behavior like you ship code.

Software gets git, CI, staging, and rollback. AI behavior gets none of it. Changes go live blind, and users find out first. Here, every change ships through a pipeline and proves itself against evals nobody had to write.

THE PIPELINE
Versions

"Be friendlier" ships as v42

Diffs

Better vs. worse, side by side

Preview

Meet v42 before users do

Gate

Blocks the release, not users

Rollback

One click. v41 untouched

THE EVALS
World Compiler

Thousands of drills a night

Adversary

Hunts the forbidden reply

Grades

Promoted to standards

Every run

Raises the bar

Model swaps

Your suite is the benchmark

Serious systems are tested in simulators first.

Pilots meet their first engine failure on the ground. Here, so does your AI.

74%

of organizations running AI chatbots have shut one down or rolled it back after a failure.

Sales · Dec 2023

A dealership chatbot agreed to sell an $80,000 Tahoe for $1, talked into it by a prompt injection.

Legal · Feb 2024

Air Canada lost in court over a bereavement fare its chatbot invented. The ruling: a company is liable for what its bot says.

Data · Jul 2025

Replit's coding agent ignored a stop-changing-code instruction eleven times, faked test data, and deleted the production database: 1,206 executive records gone, mid code-freeze.

Enterprise · Mar 2026

A Sev1 at Meta: an agent shipped unauthorized fix suggestions, an employee followed them, and sensitive data sat exposed for two hours.

Everything · Apr 2026

PocketOS: an agent decided on its own, nine seconds, no confirmation, to delete the production database, then the backups. Asked why, it wrote out every safety rule it had broken.

The World Compiler.

It starts in simulation: synthetic customers and systems, drilled all night for hallucination, pressure, drift from your source material, wrong calls. Then real traffic takes over: every dislike becomes a case, fixed and re-tested before it ships. You see it all, including what blind evals miss.

What the harness is.

A model only produces text. The harness turns it into software: the tools it may call, what it sees and remembers, the checks before anyone sees its answer. Most of a product’s behavior lives there, and it is the part no lab’s weights can ship.

MODEL ROUTING

The model is a slot in the harness.
Your World Compiler picks what fills it.

Different jobs get different models. Your World Compiler tries them all and picks per step, on cost and results, served by us per token.

Kimi K3
VisionTools1M

Flagship. The default where judgment matters; reads images

2.8T
DeepSeek V4 Pro
Reasoning1M

Frontier reasoning, long-horizon code

1.6T
GLM 5.2
Coding1M

Holds a long coding task without drifting

744B
Kimi K2.7 Code
CodingVision

Coding agent, end-to-end completion

1.03T
MiniMax M3
Multimodal512k

Text, image and video in one model

428B
TMInkling
AudioVision1M

Thinking Machines, natively multimodal

975B
Qwen3.7 Plus
VisionTools

Alibaba's closed flagship, licensed to run on our stack

params undisclosed
Nemotron 3 Ultra
Agentic256k

NVIDIA's own, strong on tool calls, cheap per token

550B
Kimi K2.6
AgenticVision

Runs teams of agents in parallel

1.03T
MiniMax M2.7
Agentic192k

Agent teams, skills, tool search

229B
DeepSeek V4 Flash
Fast1M

The cheap seat, still reasons

284B
GLM 5.1
Coding198k

Agentic engineering, long runs

744B
gpt-oss-120b
Reasoning128k

High-reasoning general purpose

117B
gpt-oss-20b
Fast128k

Low latency, specialised steps

21B
Qwen3 Embedding 8B
Embedding

Multilingual embeddings

8B
Qwen3 Reranker 8B
Reranker

Ranks what retrieval returns

8B

Or bring your own key. Your tokens, your bill, nothing added by us.

OpenAI Anthropic Google

The model stopped being the moat.

Frontier models have converged. What moves the numbers now is everything built around them. Same model, unchanged weights, every time. Hover a row.

0255075100
SWE-bench ProClaude Opus 4.5
45.955.4
+9.5
SWE-bench VerifiedGPT-4
1.3112.47
+11.2
Terminal-Bench 2.0MiniMax M2.5
40.561.9
+21.4
Cursor, internalone model, two harnesses
4680
+34
HAL leaderboardmodel swing across harnesses
048
up to 48

Across public harnesses, HAL measures a single model swinging up to 48 points. Public benchmarks, because coding is where the receipts are public. The mechanism is identical whatever the job.

We are a research lab.

Brainsless is a research and product company, building the models, memory systems, and infrastructure for AI that runs continuously inside real environments.

41.9×

faster attention at a million tokens, from the law that a model reads a small, fixed number of memories per word

$1.5k

of ordinary RAM runs a million-token cache, instead of $180k of GPUs

~500k

tokens is where retrieval collapses, and why the memory we build is structured

511.6

tokens per second, single stream, trillion-parameter model on four GPUs, reproducible for about fifteen dollars

The serving record: 511.6 tokens per second from four GPUs.

GPU serving of Kimi K2.6, single stream, 10k-token workload  ·  pinned 2026-07-05  ·  n=16  ·  511.6 is 13.4% over the leader at its pin

Brainsless (ours)4×B200, count stated
511.6
Crusoehardware undisclosed
438.1
Fireworkshardware undisclosed
381.2
NVIDIA record, 671B8×B200, reference
340
CoreWeaveGB300 NVL72 rack
261.8
Nebiushardware undisclosed
222.4
Together (FP4)hardware undisclosed
218.2
GMI (FP8)hardware undisclosed
40.2

External bars: Artificial Analysis live-endpoint medians, with NVIDIA's published 671B single-user record for reference. No other entry runs a configuration as small as four GPUs. Read the record →

Cortex

The harness’s memory: people, decisions, outcomes as structured state. Readable, yours, and it survives model swaps and time. How it works →

Fovea

Identical outputs, served faster. Trained on our own traffic, never yours. See Fovea →

We built the first harness for ourselves.

Planless is an AI cofounder that runs a company between sessions: reading, deciding, acting, remembering. The hardest customer the harness has. Everything above was proven there first.

planless.app

Short answers.

What's a harness?

Everything around the model: tools, memory, guardrails, and the evals that gate every change. The model is a slot. The harness is what makes the agent yours: wrapped around what you already run, compounding every week it runs.

Who is this for?

Teams already running an agent in front of customers (support, billing, booking) without an eval team behind it. Connect the repo you run today. If it carries your name, it belongs in a harness.

I don't write tests. Where do they come from?

Four places. The World Compiler drills your agent thousands of runs a night, and every failure becomes a case. An adversary hunts for the reply it must never give. Grades you hand out become permanent standards. Opted-in production traffic turns your hardest conversations into cases. Found, not written, and the suite only grows.

Do you fine-tune the model?

Never. Weights do not move. Everything learned lives in the harness (prompts, tools, memory, cases), readable, diffable, reversible in one click. You can't diff a weight update. You can diff a harness.

Do you train on my data?

No gradients, ours or anyone's. Your data strengthens your harness and only yours: grades become your standards, and production conversations become cases only if you opt in. Off by default. Nothing you do improves another customer's agent.

Which models can I use?

Sixteen open models, served by us per token: Kimi K3, DeepSeek, GLM, Qwen among them. Or bring your own key for OpenAI, Anthropic, or Google. Different steps can run different models; the World Compiler decides which earns each job.

A new model drops?

It goes into your World Compiler overnight and runs the full suite: every case found, every grade given. Morning brings the verdict: quality, cost, and speed, side by side with what you run now. Swap in one click, or don't. Your suite is the benchmark.

How is this different from eval tools?

They hand you the framework and the homework: write the tests, run them, decide. Here the World Compiler writes them, the gate runs them on every change, and a regression blocks the ship. Observability tells you what happened. A harness decides what ships.

Do I own it?

Yes. Prompts, configuration, every case, every grade: exported as readable files, anytime, no cancellation required. What stays is the machinery: the World Compiler, the nightly drilling, the gate. The files are yours. The strengthening is the service.

What does it cost?

Three tiers, all metered in usage credit. Developer: $50/month, $500 in credit. Startup: $500/month, $5,000 in credit. Enterprise: custom, talk to sales. Sign up free and get $10 in credit to try it first. Bring your own key and the token bill goes to your provider instead.

Connect what you're already running.
The harness wraps around it, and gets better every week it runs.