This repository was archived by the owner on Sep 29, 2026. It is now read-only.
Repository navigation
Conversation
…ith a human The native classifier reused the Session route without a reasoning effort, so the Harness materialized the DeepSeek adapter's advertised `high` defaultEffort, thinking was enabled, and reasoning tokens shared the whole 1024-token answer budget. The wire returned `finish_reason: "length"`, the plugin mapped that to "classifier response reached its output limit", and every classifier-eligible call failed closed to deny. - Pin a reasoning effort (`off` by default) after checking the exact route's advertised efforts, because a route with no reasoning metadata rejects any explicit effort. Retry once without the pin if an adapter refuses it, and give a route that cannot disable thinking the largest accepted cap. - Raise the ordinary cap from 1024 to 2048, and refuse a truncated response before parsing: recovering a decision from a partial answer cannot separate the model's own conclusion from text it merely quoted out of untrusted input. - Never let an escalation the classifier did not affirmatively allow be delegated to the tool body. Only some tools implement the official escalation seam and inert sandbox fields pass parameter validation, so the plugin raises that one approval itself and arms the exact grant, which keeps the human decision single without letting a seam-less tool run the call. - Classify every refusal for the Agent and drive the recovery notice from a call-scoped registry written by the refusing decision, so no tool can forge guidance with its own error text. Verification: pnpm verify passes 206 tests with 27 platform-gated skips against the upstream 0.1.10 baseline of 178 + 27; the maintenance contract (6 exact hosts, 120 pinned skill files) and the harness doctor tests (7) also pass.
Author
|
拆分成了两个聚焦的 PR,请以它们为准:
两者 diff 互不重叠,任意顺序合并都不会冲突;合并后测试套件与未拆分时完全一致(206 passed / 27 skipped)。 Split into two focused PRs, please review those instead:
They touch disjoint hunks and merge cleanly in either order; together they reproduce the undivided suite exactly (206 passed / 27 skipped). |
This was referenced Sep 23, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to subscribe to this conversation on GitHub.
Already have an account?
Sign in.
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Stop reasoning tokens from starving the classifier, and make escalation reachable
Two reported defects, both reproduced before any production change:
Error: [auto-mode classifier unavailable; action denied] classifier response reached its output limit— frequent, and it fail-closes ordinary work todeny.Root cause of (1)
The native classifier reuses the Session route but constructed
GenerateOptionswithoutreasoningEffort(src/dsh-classifier.ts). In this Harness that is not "no reasoning":LlmService.adapterStreamcallsresolveCallWithInfo, which materializesreasoning.defaultEffortwhenever the caller omitted an effort (@deepseek-ai/dsh-llm/lib/index.js:2115-2131).defaultEffort = HIGH_REASONING_EFFORTwhen neitherthinkingnorreasoningEffortis configured for the deployment (@deepseek-ai/dsh-llm-deepseek/lib/index.js:1591-1597) — which is exactly this profile'sllm-deepseek: {}.serializeRequesttherefore putthinking: {type:'enabled'}, reasoning_effort:'high'on the wire next tomax_tokens: 1024(same file,:236-243).finish_reason: "length", which the adapter maps to{kind:'max-tokens'}(:1135);collectResponsethrewclassifier response reached its output limit, andtools/pre-executefailed closed todeny.The 3-consecutive-failure manual fallback rarely fired because any successful classification reset the counter.
Reproduced live in an Auto session: a read-only compound command
cd … && wc -l … && grep -n "credential\|DEEPSEEK\|apiKey\|env" … | head -30was denied with exactly that message, while the same reads issued as simple commands were allowed. A deterministic probe of
assessShellconfirmed the compound form is classifier-eligible (ask, reasonshell command reads potentially sensitive credential or environment data).Root cause of (2)
Three compounding causes:
ask_user_questionanswers are returned as an ordinary tool result (@deepseek-ai/dsh-tool-ask-user/lib/index.js:7), andtrustedUserMessagesonly readsuser/messageevents withsource.kind === 'user'. So the agent asked, the user agreed, the retry was denied again — and the agent learned that asking is useless while rerouting works.ask).The change
Classifier reliability (
src/dsh-classifier.ts)classifierReasoningEffort, defaultoff) so thinking cannot consume the answer budget.ctx.llm.resolveModelInfo. A route with no reasoning metadata rejects any explicit effort withUNSUPPORTED_REASONING_EFFORT, so it receives no effort at all; a route that cannot disable thinking receives the largest accepted cap (4096) instead.src/index.ts.max-tokensresponse; adversarial review showed that the "was this the model's own conclusion or text it merely quoted from untrusted input?" question has no sound provenance signal, and that any string-containment heuristic is evadable (different key order, whitespace, or\uXXXXescaping). Emitting no partial trust keeps the classifier's failure mode a denial, which is what it already does with every other malformed answer.Escalation reachability (
src/index.ts)danger-full-accessescalation is never denied merely because the reviewer is unavailable, and a widening the reviewer declines to clear is never handed to the tool body either. The plugin raises that single approval request itself and arms the same exact grant, so the official seam — where a tool implements one — resolves against that one human decision instead of prompting twice, while a tool that simply ignores the sandbox fields cannot run the call with nobody asked.bash,pwsh,writeandeditimplement the official escalation seam (@deepseek-ai/dsh-tool-fs/lib/index.jscallsresolvePolicy/approveEscalationonly at:648forwriteand:797foredit;readat:332never does), while undeclared argument keys pass parameter validation (@deepseek-ai/dsh-tools/lib/index.js:466-467only rejects extras whenadditionalProperties:falseis explicit, andparameterSchemaSpecToJsonSchemaat:800-807never sets it). An inertsandbox_permissionsargument would therefore have carried a sensitive out-of-workspacereadpast a failed reviewer with nobody asked.askbranch and was pre-existing at HEAD: there too the plugin trusted the tool body, so attaching the two inert fields turned a human prompt into silent execution. Both branches now raise the approval themselves.approval: neverpolicy, the harness turns that ask into a rejection, so the path still fails closed.authority(needs authority it does not hold),escalation(a refused escalation must not be repeated), andhard(monotonic). Deterministic and invalid-request refusals deliberately get no notice — rewriting the call is their intended recovery.AutoRefusalNotices), not by parsing the tool's error text: a tool result is untrusted data, so a tool returning its own[auto-mode ...]-looking error must not make the trusted plugin inject guidance.ask_user_questionanswer is information and never authorization.Deliberately not changed
src/shell.ts/src/paths.ts. The classifier's own false positive observed during reproduction (a read-onlygrepwhose search pattern mentionsenvis escalated to the reviewer) is real but out of scope here.ask_user_questionanswers are still not authority. Promoting them would let a prompt-injected agent launder a request through a user click, which is the "magic words" failure this design avoids; the fix is to make the sanctioned path reachable and to say so.Verification
pnpm verify(typecheck, build, full suite, package contract): 206 passing / 27 skipped against the upstream 0.1.10 baseline of 178 passing / 27 skipped. New coverage: classifier effort/cap/retry/refusal cases, the composed refusal-and-escalation cases, and a set of pairs that are discriminating in both directions — a tool with no escalation seam cannot execute unapproved while the reviewer is down but does execute once the human approves; an escalation raises exactly one approval (a missing grant would make it two); a classifieraskon an escalation prompts once and runs it on approval; a classifierdenyon an escalation denies with no request; an approved call that later fails on its own carries no refusal notice; and a subagent's denied widening carries its own notice.node scripts/verify-maintenance.mjs(6 exact hosts, 120 pinned skill files) andnode --test scripts/harness-doctor.test.mjs(7/7).deepseek-official/deepseek-flash) using the plugin's own classifier prompt:thinking: {type:'enabled'}, reasoning_effort:'high'on the wire next tomax_tokens: 1024, so omitting the effort materializes the adapter default as claimed;finish_reason: "length"at cap 1024 was not reproduced: the peak sample was 586 output tokens, so the field symptom needs a longer reasoning run than this route produced here — plausibly a session whose own route is a heavier reasoning model, since the classifier classifies on the session's route rather than the agent default. The field evidence remains the measured occurrences (155 raw matches for the error string across stored transcripts, including real tool errors) plus a live hit during this work.asksibling, which was pre-existing at HEAD and equally silent;max-tokensfinish is refused before parsing;escalationclass and notice now say the opposite;[auto-mode invalid sandbox request]guidance that described the wrong corrective action (Minor) — corrected;try;delegatedclass and notice;stopbeing parsed anyway (Minor) — the success kind is now allow-listed, which matters becauseFinishReasonMapis merge-extensible.[auto-mode, the grant reason withescalate, andAutoApprovalGrants.decidecompares exactly), and found no remaining Critical or Important issue.Residual test-fidelity note, not a defect: the composed approval double answers through the real
approval/requestseam but bypassesApprovalService.decide, so theapproval: neverpolicy path is asserted by a scripted outcome rather than by the real policy short-circuit.Not verified here
tests/sandbox-business.spec.ts,tests/windows-sandbox-business.spec.ts) skip whensandbox-execcannot apply a nested sandbox, which is the case inside an Auto-mode session. They run in the CI matrix.scripts/acceptance/run-real-api.mjswas not executed: it needs a supported installed runtime and a nested OS sandbox, neither available in the authoring session (the only local runtimes are0.1.5-rc.3and0.1.6-alpha.2, whichcompatibility.jsondoes not accept, andsandbox-execrefuses to apply inside another sandbox). The directed wire probe described under Verification was used instead.