Skip to content
⠀⠀⠀⠀⠀⠀⠀⠀⣀⣠⣴⣶⣾⣿⣿⣧⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀
⠀⠀⠀⠀⣀⣤⣶⣿⣿⣿⣿⣿⣿⣿⣿⣿⣧⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀
⠀⣠⣴⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣧⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀
⢾⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣷⡀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀
⠀⠻⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣦⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀
⠀⠀⠹⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⡿⠿⠗⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀
⠀⠀⠀⠘⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣦⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀
⠀⠀⠀⠀⠹⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣷⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀
⠀⠀⠀⠀⢀⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀
⠀⠀⠀⣰⣿⡿⢿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀
⠀⠀⠐⠛⠁⠀⠀⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣶⣤⣀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀
⠀⠀⠀⠀⠀⠀⢸⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣷⣄⡀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀
⠀⠀⠀⠀⠀⠀⠈⠛⢉⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣷⣶⣶⣦⣤⣤⣀⣀⠀⠀⠀⠀
⠀⠀⠀⠀⠀⠀⠀⠐⠋⠁⠀⢹⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣶⣄⠀
⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠛⠛⠛⠉⠉⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣷

Lupin

Claude Code, on any model.
The gentleman router: it borrows the Claude Code harness and lends it to other models.



CI License: Apache-2.0 Node >= 20 TypeScript strict tests


Your whole setup lives in Claude Code, not in the model: MCP servers, skills, CLAUDE.md, hooks, memories, plugins. But Claude Code speaks one protocol, the Anthropic Messages API, so trying another model means changing tool and starting over.

Lupin changes the model. Nothing else moves.

Start here

Prerequisites

Required Why
Node.js 20 or newer runs the proxy and CLI
Claude Code on the PATH as claude the harness Lupin launches
One provider API key, subscription, or local runtime the model that will answer
Rust toolchain, optional only for the terminal dashboard

Lupin supports Windows, macOS and Linux. The proxy binds only to 127.0.0.1. Prompts and responses are never persisted.

Fastest path

npm install -g lupin-code
lupin                    # opens the hub; with no config, it opens setup

With the Rust sidecar on the PATH (next section) that second command is the guided add-provider screen: masked key input, a real 1-token verification, OAuth with browser polling, local runtimes with live model discovery. Without the sidecar, bare lupin starts the setup daemon and prints the authenticated curl calls that list the providers and verify your key: the first verified provider persists the config. Either way, nothing is saved before a verification succeeds, and then:

lupin run -- claude

That command starts Claude Code with its API traffic pointed at Lupin. Your existing MCP servers, skills, hooks, plugins and project instructions stay where they already are.

TUI-first setup

The optional Rust sidecar lets a new install add a provider without running init or memorizing a login command. Build it once from a clone and put the binary on the PATH:

git clone https://github.com/Fanfulla/Lupin.git
cd Lupin
cargo build --release --manifest-path tui/Cargo.toml

# macOS / Linux
cp tui/target/release/lupin-tui ~/.cargo/bin/

# Windows PowerShell
Copy-Item tui/target/release/lupin-tui.exe "$HOME/.cargo/bin/"

Then install the CLI and open the hub:

npm install -g lupin-code
lupin

With no configuration, lupin starts a temporary local bootstrap daemon and opens the add-provider screen. Nothing sensitive is written yet.

  1. Move with the arrows or j / k, then press Enter.
  2. An API-key row opens a masked field. Paste the key and press Enter.
  3. An OAuth row starts login and shows the browser URL while the TUI polls.
  4. Providers with account-suspension risk show the warning before login and require an explicit confirmation.
  5. After verification succeeds, the dashboard appears with the new profile active. The saved config keeps the same daemon identity, so the screen does not disconnect during the transition.

Local runtimes such as Ollama and LM Studio are rows in the same screen: their setup probes the live server, shows every chat model with its real window and tool support, and asks for the main and light picks before saving anything.

Choose the credential path

Every path below is a row in the TUI's add-provider screen (ADR-51: the CLI setup verbs are gone).

You have Use
A hosted-provider API key the row marked API key (masked field, 1-token verification)
ChatGPT subscription the row marked OAuth
Kimi Code subscription the row marked OAuth (offers the official-CLI import when found)
Google Code Assist its OAuth row; the suspension warning must be accepted first
GitHub Copilot its OAuth row; the suspension warning must be accepted first
Ollama, LM Studio, llama.cpp or ds4 the row marked local, with the runtime already running

Headless setup

On a machine with no TTY or no sidecar (a server over SSH, CI), the same setup lives on the daemon's control API: every route sits on 127.0.0.1 behind the local token from ~/.lupin/config.json. Add a key provider with one call:

curl -s -X POST http://127.0.0.1:3456/v1/lupin/setup-key \
  -H "authorization: Bearer $(node -p "require(require('os').homedir()+'/.lupin/config.json').localToken")" \
  -H "content-type: application/json" \
  -d '{"providerId":"openrouter","key":"sk-or-..."}'

GET /v1/lupin/providers lists the rows, POST /v1/lupin/discover-local and /v1/lupin/setup-local cover the local runtimes, POST /v1/lupin/login starts an OAuth job whose URL you open anywhere (poll GET /v1/lupin/login/:id).

Verify the first session

lupin status             # daemon, active profile and resolved models
lupin run -- claude      # start Claude Code through Lupin
lupin top                # optional live routing view in another terminal

Inside Claude Code, ask for a small tool-backed task rather than asking the model to identify itself. The model sees Claude Code's system prompt and may call itself Claude; lupin status, the TUI and GET /health are the routing truth.

If the first run fails

Symptom Check
lupin prints text instead of opening the dashboard run lupin-tui --version; the sidecar must be on the PATH and stdout must be a terminal
no config yet install the sidecar and run bare lupin in a real terminal, or use the control API (see Headless setup)
daemon not answering run lupin status, then restart with lupin stop followed by lupin run -- claude
API key rejected retry the masked field; failed verification saves nothing unless you explicitly choose save anyway
OAuth browser did not open copy the URL shown in the terminal; the CLI/TUI keeps polling
port 3456 already in use run lupin status; do not kill an unrelated process until you identify it
lupin update from 0.2.4 fails with Windows EBUSY close Lupin and Claude Code, reboot Windows, then install 0.2.5 or newer before reopening Lupin; use the recovery command below if npm was damaged too
Windows build cannot find link.exe or kernel32.lib install MSVC Build Tools and the Windows SDK, then use tui/build-msvc.bat --release

PowerShell recovery after the 0.2.4 EBUSY failure:

$nodeDir = Split-Path (Get-Command node).Source
& "$nodeDir\node.exe" "$nodeDir\node_modules\npm\bin\npm-cli.js" install --global lupin-code@latest

Config, logs and credentials live under ~/.lupin by default. LUPIN_DIR moves that whole directory. API keys and OAuth tokens live in the OS keychain when available, otherwise in a mode-600 credentials file. They never enter config.json or logs.

Switch model with the session open, no restart:

lupin use glm                # you are on GLM-5.2 from the next request
lupin use gpt --bg kimi      # main on GPT, haiku-tier traffic on Kimi
lupin go kimi-sub -- claude  # switch and launch, one gesture

Or without leaving Claude Code at all: open /model and pick a row that reads switch Lupin profile: <name>. Lupin publishes one per profile, and picking one moves the active profile from inside the session, with the conversation intact. That is the answer to running out of quota mid-task, and it is exactly how this was verified: Kimi answered usage limit for this billing cycle, the picker moved the session to ChatGPT, and the next request was served.

One honest limit of the picker

The client re-sends the picked id on every later turn, so Lupin acts on it only when it changes. Consequence: pick switch: B from the picker, then switch to A from the CLI or the TUI, and the picker cannot bring you back to B, because re-picking the row it still shows as selected sends the same id and no gesture is seen. Any other surface works (lupin use B, the TUI, or picking a different row first). The alternative was letting every turn of an old session drag the active profile back, which would mean no other surface could ever hold a switch.

Not one more router: the honest one

Everybody translates formats. Nobody tells you whether the model survives the harness.

That is the hard part. A model can speak the protocol perfectly and still fail every task, because Claude Code demands exact-match edits, a tool loop that closes, and roughly 46,000 tokens of prompt before the first word. So Lupin ships the question as a command:

lupin doctor kimi-sub
lupin doctor scoring the ChatGPT subscription 10/10
A real run against the ChatGPT subscription. Only the waiting is edited: the session took 116 seconds and is held for two.

lupin doctor runs a real headless Claude Code session against a dedicated server on an ephemeral port, and scores it from artefacts on disk (files really edited, scripts that really run), never from what the model claims it did. Six checks, threshold 7, no hidden retries.

Three things it refuses to do, each of them learned the hard way:

  • It will not grade a session that never reached the model. Claude Code reports success when its own loop ends cleanly, even when every request died on a protocol error. Where another tool would print 1/10 and blame the model, the doctor prints notRun and the cause (ADR-23).
  • It will not hide its own help. A 10/10 earned because the proxy repaired a broken tool call is different information from a clean 10/10, and the verdict says which one you got.
  • It will not invent a number. No score is reported for a provider nobody has run.
What the doctor has actually measured (dates included, because a score without one is a rumour)
Profile Score When Notes
kimi-sub (Kimi Code subscription) 10/10 2026-07-28 reproducibility measured separately on 2026-07-19: 10, 10, 10
kimi (Moonshot API key) 10/10 2026-07-19 95s, cache_control accepted
openai-sub (ChatGPT subscription) 10/10 2026-08-05 116s over 15 requests, 49% of input served from cache (the run in the GIF above)
gemini-sub (Google Code Assist) 8/10 2026-07-29 free tier: the run ended on 429 No capacity, which is a tier limit, not a translation defect
lmstudio + gemma-4-12b 0/10 2026-07-19 honest: the harness floor and this GPU cannot be satisfied together
copilot-sub not run works live, but a full doctor run would spend most of a free plan's monthly allowance

Providers

Four lanes, picked by Lupin, never by you. Passthrough first: when a provider already speaks Anthropic, nothing is translated at all.

Provider Lane Credential State
Kimi / Moonshot passthrough API key or subscription verified live, doctor 10/10
ChatGPT subscription responses Sign in with ChatGPT (PKCE) verified live, doctor 10/10
Gemini Code Assist subscription codeassist Sign in with Google (PKCE) verified live, doctor 8/10
GitHub Copilot subscription translate GitHub device flow verified live 2026-08-05. On the free plan expect about 50 chat requests a month, which one real session can spend
DeepSeek, Z.AI / GLM passthrough API key implemented, not scored
OpenRouter passthrough or translate API key 344 models, 255 with tool calling
OpenAI, Gemini (pay per token) translate API key implemented, not scored
Ollama, LM Studio, ds4-server passthrough none local, native Anthropic endpoint
llama.cpp server translate none local

344 models on OpenRouter, and Claude Code can natively use zero of them, because OpenRouter's Anthropic-compatible endpoint only accepts Anthropic models. Through Lupin the 255 with tool calling become usable. The other 74 connect and stay a chat, which Lupin says out loud instead of letting you find out mid-task.

How the routing works
Claude Code ──ANTHROPIC_BASE_URL──▶ Lupin (127.0.0.1)
                                      │
            ┌───────────────┬─────────┴─────────┬────────────────┐
            ▼               ▼                   ▼                ▼
      [passthrough]    [translate]        [responses]      [codeassist]
      Kimi, DeepSeek   OpenAI, Gemini,    ChatGPT          Gemini Code
      Z.AI, Ollama     llama.cpp,         subscription     Assist
      LM Studio, ds4   OpenRouter,        (WHAM)           subscription
                       Copilot
  • Passthrough rewrites the URL, the auth header and the model name. Nothing else is touched, so the provider's prompt cache keeps hitting: on a local runtime that is the difference between a turn measured in seconds and one measured in minutes.
  • Translate maps requests, responses and SSE streaming, tool calling included, with MCP names longer than 64 characters rewritten through a deterministic hash.
  • responses and codeassist exist because a ChatGPT or Google OAuth token does not spend on those providers' public APIs at all. Each subscription has its own private protocol, so each got its own translator, built from real captured traffic.

Three model slots (opus, sonnet, haiku) map to whatever you point them at, per profile; --bg sends the haiku-tier traffic somewhere cheaper, and the agents table (SPEC-PROVIDERS §4decies) routes each subagent type to its own model or provider.

Where routers really break

A provider can honour the protocol to the letter and still hand you garbage inside the content: reasoning wrapped in <think>, a tool call the server never turned into tool_calls[], special tokens leaking into the text. The model "calls" the tool, nobody runs it, and the agentic loop dies quietly.

In a sample of 382 claude-code-router issues, over 29% land here, re-fixed provider by provider instead of once. Lupin treats it as one problem, with one engine.

The four rules that engine follows
  • One engine for both paths. The same normalizer serves the non-streaming mapper and the SSE translator, and a test feeds it the same input character by character and as a single block, demanding identical results. A normalization that behaves differently while streaming is exactly the bug this makes impossible.
  • Verified markers, never guessed. Eight model families (Qwen3, Qwen3-Coder, GLM, DeepSeek, Kimi K2, Mistral, Llama, GPT-OSS Harmony) checked against the official chat templates and the parsers of vLLM, llama.cpp and SGLang, then put through an adversarial pass that threw out the invented ones. GLM reuses Qwen's <tool_call> with a different payload; DeepSeek uses fullwidth vertical bars where Kimi uses ASCII. A wrong character does not fail loudly, it simply never matches.
  • Reasoning is not lost. reasoning_content becomes a thinking block. On gemma-4-12b the content field arrives empty and the whole answer is in there: drop it and Claude Code receives nothing.
  • Never silently. Every normalization that fires lands in the log and in the doctor verdict.

The terminal, done properly

A bare lupin opens the hub. With the optional Rust sidecar on your PATH it is a live dashboard; without it, a status summary and the next step.

The sidecar is not on npm and never will be: it stays out of the JavaScript runtime by design. Build it from a clone, once, then launch it with the bare command:

cargo build --release --manifest-path tui/Cargo.toml
cp tui/target/release/lupin-tui ~/.cargo/bin/      # anywhere on PATH will do

lupin        # the hub finds the sidecar and opens the dashboard
lupin-tui    # or run it directly

That build is a one-time cost: from then on lupin update rebuilds the sidecar to the matching version on every package update, from the sources the package itself ships (it needs the Rust toolchain; without one it prints the manual command instead).

Keys: 1-9 switch profile, arrows and Enter do the same on the highlighted row, d runs the doctor on the highlighted profile, t tries a model (type, search the provider's catalogue or paste an id, then land it on the whole profile in one gesture, single slots excludable), m aims its slots (opus, sonnet and haiku edited in place, written as given and never checked; with a catalogue the focused field suggests while you type and Tab completes), : opens the command palette, o edits the failover order, a opens agents mode, r refreshes now, q quits.

Providers that publish a model list (OpenRouter first) feed those inputs live: each row shows the served context window, tool support and price, the picked window lands in the profile's contextWindows in the same write, and an id outside the list is still written as given, with an advisory instead of a refusal. Pasting works everywhere a model id is typed.

Agents mode is the subagent mixer on screen: it lists every agent route plus the conventional subagents row (shown even before it exists, so the first gesture is obvious), and on the selected row 1-9 aims it at that profile, m aims it at a model id (catalogue-assisted), n names a brand-new route, x clears it, Enter applies the whole table atomically through the control API, Esc throws the edit away. After applying a named route the TUI offers to wire the agent file's model: line for you (explicit y, skippable, the route stays saved either way). The daemon writes the config and hot-reloads it, so a live Claude Code session picks the new routing up on its next request.

The CLI exposes the same advanced routing without opening the dashboard:

lupin agents set subagents --profile ollama-qwen   # every subagent on the local model
lupin agents set explore --profile kimi --wire     # wire one named agent to Kimi
lupin agents                                       # inspect routes and model ids

lupin agents prints claude-lupin-agent:<name> for each route. Use that id in an agent definition's model: field, or pass --wire to update that one field explicitly. The blanket subagents route is carried by CLAUDE_CODE_SUBAGENT_MODEL, which lupin run sets for you. Each routed request remains visible as agent:<name> in lupin top and the log.

The doctor takes minutes, so it runs as a child process and its output streams into a panel while the dashboard keeps refreshing underneath. Provider setup is native to the TUI, whole (0.3.0, ADR-51): API-key rows verify before storing, offer the economy preset where one exists, and turn a provider rejection into an explicit save-anyway choice; OAuth rows poll the browser flow, confirm suspension risks first, offer the official-CLI credential import when one is found, and take an optional account label so a second account gets its own profile; local rows (Ollama, LM Studio, llama.cpp, ds4) probe the live server and list every chat model with its real window, tool support and a context-too-small verdict before you pick the main and light models. Rows that already have a profile say configured, and the failover offer arrives only after a setup succeeded, never before the provider answered. Opening the hub also makes sure the daemon is running, so the dashboard never starts dead. The palette runs doctor, usage, list, status and stop. Its only shell-only row is run, because Claude Code needs to own the terminal.

Give it 32 rows or more and it draws the portrait full size; below that it keeps every fact and shrinks the art. It needs a real terminal: it takes over the screen, so it will not do anything useful inside another tool's output pane.

⣀⣤⣶⣾⣿⣷⣄⠀⠀⠀⠀⠀⠀⠀⠀⠀  L U P I N  v0.3.2   the gentleman router
⠈⢿⣿⣿⣿⣿⣿⣷⣖⠀⠀⠀⠀⠀⠀⠀  daemon up   127.0.0.1:3456
⠀⠴⢿⣿⣿⣿⣿⣿⣿⣀⠀⠀⠀⠀⠀⠀  active: kimi-sub  ->  k3
⠀⠀⠈⠛⢿⠿⣿⣿⣿⣿⣿⣶⣶⣶⣤⣄

Profiles with 1-9 hotkeys, the routing truth per slot, the request tail with every marker (routed, agent, failedOver, tierDowngrade, dialect), and a status line that says in words what just happened. lupin top gives the same truths with no sidecar at all. Details in docs/TUI.md.

There is also a statusline for Claude Code itself, because through a proxy the model introduces itself as the Claude of the system prompt (the UI knows nothing about the mapping, ADR-3), so asking it who it is proves nothing. The truth lives in GET /health, and the statusline shows it as ⇄ profile→model.

Statusline install and every segment it draws

Opt-in, always: Lupin never writes your settings.json (ADR-11). Copy examples/statusline.ps1 (Windows) or examples/statusline.sh (macOS and Linux, needs jq) into ~/.claude/ and register it:

"statusLine": { "type": "command", "command": "powershell -ExecutionPolicy Bypass -File \"C:\\Users\\<you>\\.claude\\statusline.ps1\"" }
Segment Example What it says
Skill flag [CAVEMAN] active mode or skill (flag file), specific to a personal setup
Model Fable 5⚡ the model requested by Claude Code. Through a proxy this is the slot name, not the real model
Lupin routing ⇄ kimi-sub→k3 active profile and the real model of the opus slot, from /health with a 10s cache. Daemon down shows a red OFFLINE
Repo Lupin@main +3/-1 folder, git branch, uncommitted lines
Context ctx: 67k/1M (7%) uses total_input_tokens, cache reads included, which is what makes the percentage honest
Effort effort: xhigh reasoning effort, coloured by rising cost
Thinking ✦think extended thinking active
Cost $4.20 suppressed through Lupin: Claude Code prices Anthropic models, so on another provider it would be fiction
PR PR#12✓ PR state of the branch
Quota 5h: 30% reset 14:00 Claude subscription limits, which through Lupin disappear on their own
Update ↑v2.1.216 a newer Claude Code on npm

Three lines maximum: when space runs out segments drop in order (update, PR, extras, 7d, 5h). The mandatory ones always stay.

Questions people actually ask

Can I use my ChatGPT subscription with Claude Code?

Yes. The ChatGPT OAuth row in the hub uses the sanctioned Sign in with ChatGPT flow. The token does not spend on the public OpenAI API, so Lupin talks to the same protocol the official client uses, and lupin doctor openai-sub scores 10/10 on a real session.

Can I run Claude Code on Gemini for free?

Yes, with a caveat worth reading. The Google OAuth row in the hub reaches Google Code Assist, whose free tier answers on the flash models and returns 429 on the pro ones. Lupin serves you rather than refusing, and logs every substitution as tierDowngrade so you always know which model answered. Two honest warnings: Google collects prompts and code on the free tier with human reviewers able to read them, and Google has suspended accounts for third-party OAuth, which is why the login blocks on an explicit risk confirmation first.

Does my claude-mem / MCP / skills setup keep working?

Everything local does: native tools, local and project MCP servers, plugins, skills, hooks, CLAUDE.md, memory, subagents. Verified with real sessions. What breaks is tied to the claude.ai account, and it breaks with any proxy, not just this one: see the section below.

What does it cost me to run?

Nothing beyond the provider. Lupin is a local process on 127.0.0.1, it has no backend, it uploads nothing, and it never persists prompts or responses. lupin usage aggregates your own log offline, and it sees subagent traffic that the Claude Code transcript does not: 332 requests against 113 visible turns, measured.

Why is Anthropic not in the provider list?

Because Claude Code already runs Claude natively and does it better than a proxy would (ADR-18). Lupin exists to reach the models Claude Code cannot.

What you lose with any proxy

Claude Code ties some features to the claude.ai login and to api.anthropic.com. Point ANTHROPIC_BASE_URL anywhere else (Lupin, claude-code-router, LiteLLM, any gateway) and Claude Code itself disables them. There is no hybrid mode.

  • Remote Control (driving the session from claude.ai or your phone), since v2.1.196.
  • MCP connectors hosted on claude.ai, voice dictation, and the cloud surfaces (web, mobile, Slack, routines, ultrareview).

One cosmetic quirk to expect: resuming a session may print Session model k3 could not be restored. Claude Code persisted the provider's own model name and fails to find it in its catalogue. The routing stays correct, you only lose the persistence of the /model selection across restarts. Lupin does not rewrite the response to hide this, because byte-faithful passthrough is the entire point of passthrough (ADR-7).

Local models, zero keys

Pick ollama, lmstudio, llamacpp or ds4 in the wizard: no key to paste, the models are read from your own server with their real context windows, and you get a warning about the ones that do not declare tool support, since Claude Code cannot take a single step without them.

That distinction is not pedantry. gemma-4-12b declares a 262,144 token window and runs with 8,192: a factor of 32. Lupin always prefers the loaded window over the declared maximum, and marks which one it got.

Everything else

Command reference
Command Does
lupin the hub: TUI when the sidecar is installed, else status and next steps
init wizard: provider, key (never echoed), a real connectivity test
login <provider> / logout OAuth, with --account <label> for a second account on the same provider
use <profile> [--bg <p>] hot switch, no restart: the open session moves on its next request
go [profile] -- <cmd> switch and run in one step
run -- <cmd> start the daemon if needed and run with the env pointed at Lupin
resume [profile] continue this directory's last session on another provider
doctor [profile] the real headless session, scored on disk artefacts
use <profile> --opus <model> aim a slot by hand, for profiles whose models come from the account
agents set <name> --profile <p> [--wire] per-subagent routes; --wire writes the agent file's model: line for you
update update the npm package and rebuild the TUI sidecar if you have one
list / status / stop / logs -f the plain truths
top live console, no sidecar needed
usage [--days N] tokens really served, aggregated from your local log

Every command behaves identically on Windows PowerShell, cmd, and any POSIX shell. lupin run spawns Claude Code with no shell in between, so your arguments arrive byte for byte (ADR-29).

Documentation index
File Content
docs/NEXT-STEPS.md Start here: current state, how to verify it, what to do next
DESIGN.md Vision, prior art, positioning, risks
docs/DECISIONS.md ADR log: every decision, the why, the rejected alternatives
docs/SPEC-TRANSLATION.md Translation core: mapping, SSE, errors, acceptance fixtures
docs/SPEC-PROVIDERS.md Provider registry, profiles, slot mapping, quirks
docs/SPEC-CLI.md CLI, doctor, security, UX
docs/ROADMAP.md Milestones, verification criteria, next steps
docs/ARCHITECTURE.md Repo layout, dependency rules (a pure core)
docs/TESTING.md Fixtures from real output, test levels
docs/TUI.md The terminal hub: install, keys, panels, troubleshooting
docs/COMPETITIVE.md Competitive analysis: white space, steal candidates
docs/DESIGN-OAUTH.md Pluggable credential source, device flow
docs/DESIGN-OAUTH-PKCE-TUI.md OAuth PKCE, the control API, the Rust sidecar
docs/DESIGN-TRANSLATORS-DEDICATED.md The two subscription translators

Principles

  1. Fixture first. The fixture, recorded from real provider output, comes before the code. Real dialects are stranger than you would guess: keep-alive SSE comments, repeated finish_reason, usage arriving after the stream ended, errors delivered as a data frame.
  2. Centralized quirks. Never if (provider === x) scattered around. Flags in one registry, one implementation each.
  3. Privacy. Prompts and responses are never persisted. Keys live in the OS keychain or a 600 file, never in the config and never in the logs. It binds to 127.0.0.1 only.
  4. Zero side effects. Lupin never touches ~/.claude/settings.json. Uninstalling means stopping using it.
  5. No invented numbers. A context window enters the defaults only when the vendor publishes the exact figure. DeepSeek and Gemini write "1M" without saying whether that is 1000 or 1024, so they stay without one: a route that never fires beats a route that fires on the wrong number.
  6. Disciplined scope. Lupin stays a proxy. No Electron app, no web dashboard, no relay bot. The biggest competitor accumulated 853 open issues while adding surface; the answer here is not to add it.

The portrait is Arsene Lupin as Leo Fontan drew him in 1908 for Arsene Lupin contre Herlock Sholmes. Public domain, like the books.

Apache 2.0

About

Run Claude Code on any LLM provider without losing your setup, local-first proxy with verified profiles, OAuth device-flow login, failover, content-aware routing and an honest per-model compatibility doctor

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

34 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages