RetrospectiveWhen Code Can't Leave the Building: Janus, a Local LLM Agent IDE
· Janus
- #local-llm
- #harness
- #agent
- #mlx
I once assisted at an AX (AI transformation) training at a manufacturing company. Security mattered more than anything there. Code couldn't leave the internal network, so external AI services were hard to use.
A developer I know who works in finance was in a similar spot. They use AI, but with a lot of restrictions.
Around then Qwen3.8 came out and got a good reception. I wondered whether a coding agent could run locally, and started building Janus as an experiment. It's an agentic coding IDE built on a local LLM.
Starting point: a weaker local model
A local model can't match Claude or Codex. Janus's default model is Qwen3.8 27B, quantized to 4-bit and run with MLX (Apple's machine learning framework for Apple Silicon). On a 48GB Mac the model alone takes 16GB, and when several requests run at once, each one keeps its own KV cache (memory that stores earlier token computations).
Since the model couldn't change, I looked at how far the harness around it (the part that runs, verifies and budgets the work) could raise the results.
27B
Local model
Qwen3.8 · MLX 4-bit
44/45
Fixed tasks passed
3 tasks × 3 delegation modes × 5 runs
0
Code sent off the machine
on the local path
Deciding "done" outside the model
A weaker model sometimes says it's finished without running the tests, or stops with the tests failing. So in Janus, the end of the model's answer doesn't count as the end of the task.
Agent run
A finished answer is not done yetGit diff
Changes recomputed from the real repoVerify
The harness runs tests, lint, acceptanceReview
A person checks the diff and resultsCommit
Push only if the recorded commit is HEAD
Run finished, verified, review accepted and committed are recorded as separate states. Changes are computed from the repository's own Git diff rather than a separate snapshot, and when the revision (an ID computed from the whole change set) changes, earlier reviews and verification have to be redone. Nothing can be committed without passing verification.
Delegation with a budget
I expected that splitting large tasks across sub-agents (workers) would help a weaker model. Measurements on the real model didn't match that.
Autonomous delegation
All 15 runs passed, but at 72.5 seconds on average, 63% slower. On the multi-file task it created unnecessary workers in all 5 runs.
One worker for an investigation task
All 5 runs failed. The worker used up its token limit (8,192) each time and kept respawning the same sub-task: 4 runs timed out and 1 failed verification.
So delegation is now off by default and allowed, with a budget, only when the user asks for it.
- Worker count and token, time and step budgets are capped, and a role can be respawned only for one initial try and two corrections.
- When a worker runs out of budget, its changes and partial results go to the parent agent so the same work isn't assigned again.
- Only one role edits files; research and verification roles get read tools only. Calling a write tool that wasn't given is rejected at execution.
- Retries depend on the failure. If verification fails, a sub-agent that can only read finds the cause and the sub-agent that edits files fixes it; budget, timeout and tool errors each set up the next attempt differently.
Saving local resources
Locally, generating less means finishing sooner, so Janus avoids resending or regenerating the same content.
- Older conversation is summarized while keeping tool calls paired with their results. In a test that simulates long sessions, input dropped by more than 40%.
- The unchanging prefix (system instructions and summary) is connected to the model server's prefix cache (which reuses the computation for an identical beginning). On the second turn of a session, 36 of 57 prompt tokens were served from the cache.
- Concurrent generation went from 5 to 3. Each request keeps its own KV cache and there is one GPU, so going past 2 or 3 brought no gain, and in practice concurrency never went above 3.
Judging improvements by measurement only
Each harness change was compared against the current setup by running a fixed set of tasks (each with a test that decides whether it is done) several times, with the same model and delegation mode. If the pass rate dropped, the change didn't count as an improvement even if time and tokens went down. A change to the scheduler, which decides the order of work, dropped the pass rate from 44/45 to 40/45, so it wasn't applied.
The single-worker mode kept the pass rate while using less time and fewer tokens, so it was applied.
14/15
Acceptance passed
same as before
88.1s
Wall time
-19.8%
10,993
Prompt tokens
-26.0%
777
Completion tokens
-36.7%
Looking back: what the harness covered and what it didn't
Deciding completion by verification, capping delegation and checking changes by measurement let the 27B model pass 44 of 45 tasks.
The model's own limits remain. Deep investigation and decisions across many files still depend on the model, and the difference is obvious next to Claude or Codex. That's why Janus also lets you use an already signed-in Claude Code or Codex CLI instead of the local model, with the documentation spelling out how control differs on each path.