feat(agent-core-v2): add runtime subagent model failover - #2344
Draft
xiayh0107 wants to merge 1 commit into
Draft
Conversation
🦋 Changeset detectedLatest commit: db38bab The changes in this PR will be included in the next version bump. This PR includes changesets to release 1 package
Not sure what this means? Click here to learn what changesets are. Click here if you're a maintainer who wants to add another changeset to this PR |
Author
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.


Related Issue
No linked issue. This draft PR is intended to align on the runtime failover design before it is considered ready for review.
Problem
Subagents can select a secondary model when they are created, but an already-running subagent cannot recover by switching models when its provider remains unavailable. Existing step retries only replay the same driver with the same model, error recovery stops after the first matching handler declines, and each turn keeps a frozen request configuration.
As a result, retry exhaustion or a structured provider quota error fails the turn even when an ordered fallback model is configured. Retried streams can also leave abandoned partial text or tool frames in transcript and Web projections.
What changed
[subagent_failover]configuration andKIMI_CODE_EXPERIMENTAL_MODEL_FAILOVERflag with ordered model/effort bindings, trigger selection, and a per-turn switch limit.model.failoverwire records andturn.step.failoverevents with model, provider, effort, reason, and switch-budget metadata.The first version intentionally does not change the v1 engine or interactive TUI.
Documentation note: the
gen-docsskill was run, but its prerequisite check requiresdocs/scripts/sync-changelog.mjsand a rootCHANGELOG.md; neither exists in this repository. The skill therefore stopped as instructed. Generated agent-core-v2 manifests are included.Validation:
pnpm typecheckpnpm lint(0 errors; repository baseline warnings remain)pnpm --filter @moonshot-ai/agent-core-v2 lint:domainpnpm exec changeset statusChecklist
gen-changesetsskill, or this PR needs no changeset.gen-docsskill, or this PR needs no doc update. See the prerequisite note above.