Repository navigation
Add the roboflow-workflow-evals skill - #51
Closed
joaomarcoscrs wants to merge 4 commits into
Closed
joaomarcoscrs wants to merge 4 commits into
joaomarcoscrs wants to merge 4 commits into
Conversation
Teach agents when and how to evaluate a Roboflow Workflow: Specs, Eval Datasets, Cases, bindings, Runs, results, and cleanup. The reference folder is a generated snapshot of the Workflow Evals engine release. CI validates every skill folder, keeps generated references to engine-sync pull requests, and a scheduled job marks those drafts ready once production serves the engine version.
Older drafts that production can also serve are reported as superseded instead of marked ready, and a draft whose engine version cannot be read no longer stops the others. Note per-provider errors in embeddings reads.
Compare every open engine-sync pull request and the reference on the default branch, promote only the newest draft production can serve, and close older drafts with the reason, so a stale snapshot is never marked ready on a later run.
Leave a draft open while only an unmerged newer sync outranks it, and promote or close nothing in a run where any open sync pull request's engine version cannot be read.
Contributor
Author
|
[Jarbas Local João] — APPROVED at Gist: https://gist.github.com/joaomarcoscrs/176b4b8922f06dbada159a59d42e0ed2 |
Contributor
Author
|
Closing: we are keeping this change to the Workflow Evals bug fixes (roboflow/roboflow#16897 and roboflow/roboflow-mcp#206) and not changing how skills and reference material are organized for now. The branch stays available if we pick this up later. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Adds
roboflow-workflow-evals, the skill that teaches agents when and how to evaluate a Roboflow Workflow against ground truth. Workflow Evals goes GA this week, and until now agents only learned it from MCP tool descriptions.The skill
SKILL.mdauthoring.mdlifecycle.mdtroubleshooting.mdreference/create-evalprocedure, and examples for one engine release, listed inreference/manifest.json.The guidance comes from an end-to-end run of every Workflow Evals MCP tool against production, plus the platform and engine code. Behavior that differs by engine version (evaluator types, schemas) is read from
reference/when versions match and fromworkflow_evals_reference_getwhen they do not. The skill tells agents to checkworkflow_evals_capabilitiesfirst.roboflow-inferenceandroboflow-training-and-evaluationnow point here for Workflow-level evaluation. The training skill stays under the 20,000-character limit (19,933).Keeping it current
reference/is not edited by hand. On each engine release, roboflow/workflow-evals regenerates it withnpm run agent:export-public, lets an agent update the guides, and opens a draft pull request here labeledengine-sync. The private release notes go to that agent only, never into the public pull request.New CI in this repository:
SKILL.md, 20 companion files), relative links, nestedSKILL.mdfiles, and that a generatedreference/declares exactly the files it ships. A second job fails a pull request that changesskills/*/reference/*without theengine-synclabel. This PR carries the label because it adds the first snapshot.engine-syncdraft ready once production'sworkflow_evals_capabilitiesreports that engine version, so the skill never documents an engine the API does not run yet. A draft whose engine is not newer than the reference onmainwould roll the reference back, so it is closed with the reason. A draft that only an unmerged newer sync outranks waits. If any open sync pull request's version cannot be read, the run promotes and closes nothing.The snapshot here is engine 0.5.0. Production runs an older engine today, and the version check covers that gap.
Setup needed after merge
ROBOFLOW_API_KEYand variableROBOFLOW_WORKSPACEfor Promote engine sync. Any workspace works; the job only readscapabilities. Without them the job logs a warning and does nothing.Cowork
The skill joins the Cowork package (11 of 20 skills) and the package version goes to 1.1.0.
python3 cowork/build.pybuilds and validates it locally.Related
workflow_evals_reference_getand adds the fixes this skill relies on. Dependabot then bumps the MCP's skills submodule, which serves this skill asroboflow://skills/roboflow-workflow-evals/SKILL.Checks run
python3 .github/scripts/validate_skills.py: 11 skills valid.python3 -m unittest discover -s .github/scripts: 17 tests.python3 -m unittest discover -s cowork -p "test_*.py"andpython3 cowork/build.py.