Repository navigation
Add skills as slash commands - #711
Draft
hanna-paasivirta wants to merge 14 commits into
Draft
hanna-paasivirta wants to merge 14 commits into
hanna-paasivirta wants to merge 14 commits into
Conversation
# Conflicts: # services/global_chat/PAYLOAD_SPEC.md # services/global_chat/global_chat.py # services/global_chat/router.py # services/global_chat/tests/unit/test_planner.py
2 of 7 tasks
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Short Description
Adds skills to global chat:
/design,/qaand/diagnose. A user can run one as a slash command, or the planner can load one when the task matches. Skills can also be written for the job code agent, which the planner can hand over by name or the job agent can load itself.Decisions
skill. Apollo never parses commands out of the message.load_skilltool. The planner loads a skill when the task matches, not only when the user's wording does, so a failed run attached to "hi" is diagnosed with/diagnose.services/global_chat/skills/<name>/SKILL.md). OnlySKILL.mdis read for now, but the folder is there for reference files later.metadata: {agent: job_code}in its frontmatter). The planner can name it on a job agent call, so fixed instructions reach the job agent word for word instead of being rewritten on every call, and the job agent can also load it itself, e.g. when a user asks straight from a step page whether the step is ready to go live. Each agent sees only its own skills./qasplits this way: its per-step review lives inqa-code, which a future skill (say/migrate) can reuse. The planner reads the whole workflow first and settles the contracts between steps, then has every step reviewed and fixed in one parallel pass, then checks the result. The planner's message carries only what the job agent can't know (contracts, that other steps are being fixed at the same time, and what the user asked, such as "leave comments alone") and overrides the skill where they conflict.Implementation Details
skill: {name}on the turn the user types the slash command. That turn skips the router and goes to the planner, with the skill's instructions ahead of the message.load_skilltool, so it can pick up a skill when the user didn't type the command. Its prompt now mentions skills alongside tools.history, so a skill like/designkeeps working over several turns. While a skill is in the history, every turn goes to the planner, because the other agents never saw it.[pg:...]page prefix in history, like the other routes, so all routes return history in the same shape.call_job_code_agenthas an optionalskillparameter, listing only job-agent skills. Code puts the named skill ahead of the planner's message. Job-agent skills can't be invoked by slash command or loaded by the planner.skillslist in its payload, which global_chat sends on both the planner and the router's direct route (leaving out a skill the planner already attached). job_chat offers them through its ownload_skilltool, which continues its tool loop the wayinspect_job_codedoes, and its per-turn reminder mentions skills when it has any.meta.skillslists every skill a turn used: invoked, loaded by the planner, attached to a job agent call, or loaded by the job agent itself. Each becomes askill:<name>tag on the Langfuse trace (checked on a real traced turn), so traces can be filtered by skill.global_chat/tests/acceptance/skills/for the core skill behaviours, rather than one per skill: the job agent reviewing a step it's handed as done, the same step with one specific change (no full review),/designstill consulting on its second turn, and/qaend to end (cross-step fix, comments left alone, adaptor pinned to the tested version). All four pass. The acceptance output now shows the skills used next to the agent path.Required Lightning changes
These are in OpenFn/lightning#5236.
historyApollo returns exactly as received (newapollo_historycolumn), and send it back ashistoryon the next request. Don't rebuild it from messages, or the skill instructions and page prefixes are lost. Old sessions can still fall back to rebuilding.historyfrom the finalcompleteevent ofworkflow_chatandglobal_chattoo, not onlyjob_chat. It's there for all three.skill: {name}only on the turn the user invokes the slash command, not on later turns.designto Lightning's skill list. Don't addqa-code; it's for the job code agent, not users.