Repository navigation
fix(dataset): name the malformed YOLO line instead of failing on its shape - #2665
Conversation
_parse_box built xyxy as x_center +/- width/2 without checking the sign, so a label such as '0 0.5 0.5 -0.4 0.4' produced [70, 30, 30, 70]: x_min past x_max. Nothing downstream expects that ordering and nothing reported it. Detections.area came back -1600, box_iou_batch scored two boxes covering the identical region at 0.0 instead of 1.0, and with_nms kept both of a pair it should have suppressed, so a corrupt label silently skewed evaluation. Reject a negative extent on load, matching how this parser already rejects a class id that is not a whole number and a non-numeric trailing token. Detections does not enforce x_min <= x_max, so object_to_yolo could write the negative extent the loader now refuses, leaving supervision unable to read a dataset it had just written. Order the corners before measuring the width and height so both ends agree; the exported box describes the same rectangle and normal boxes are byte-identical. A zero extent is degenerate but still ordered, so it keeps loading. The polygon path is unaffected: polygon_to_xyxy takes min and max. NaN and inf extents were already rejected by Detections' finite-value check.
…box-extent # Conflicts: # docs/changelog.md
…shape
yolo_annotations_to_detections appended a class id for every line but only
appended a box when the line matched a known shape, with no else branch. A
line of one to four tokens therefore left class_id_list one entry longer than
relative_xyxy_list, and the mismatch surfaced far from its cause:
'0 0.5 0.5 0.2' -> operands could not be broadcast together
with shapes (0,) (4,)
'' -> IndexError: list index out of range
one good line, one truncated -> class_id must be a 1D np.ndarray with
shape (1,), but got shape (2,)
The last is what a user sees when one row of an otherwise good file is
damaged, and it blames class_id rather than the short line.
is_obb reads nine-token four-corner lines only, and any other count died
inside _parse_polygon's reshape with 'cannot reshape array of size 5 into
shape (2)'.
Check the token count before parsing and name the offending line, matching
how this parser already rejects a class id that is not a whole number and a
non-numeric trailing token.
Codecov Report✅ All modified and coverable lines are covered by tests. Additional details and impacted files@@ Coverage Diff @@
## develop #2665 +/- ##
=======================================
Coverage 92% 92%
=======================================
Files 78 78
Lines 11900 11912 +12
=======================================
+ Hits 11004 11018 +14
+ Misses 896 894 -2 🚀 New features to boost your workflow:
|
There was a problem hiding this comment.
Copilot review overview
🟡 Changes recommended
Regression tests do not yet verify the central guarantee that errors identify the malformed annotation row.
Review effort: Balanced
Findings: 1
Open (1)
What changed in this PR
Improves YOLO dataset-loading diagnostics, alongside the stacked box-extent fixes.
Changes:
- Reject malformed token counts with errors identifying the annotation line.
- Reject negative box extents and normalize exported box corners.
- Add regression tests and changelog entries.
Quality: Code 4/5 · Testing 3/5 · Documentation 4/5
Risk: 2/5 — localized parser and export changes.
| File | Description |
|---|---|
| tests/dataset/formats/test_yolo.py | Tests malformed lines, box extents, and export ordering. |
| src/supervision/dataset/formats/yolo.py | Validates annotation structure and normalizes box extents. |
| docs/changelog.md | Documents diagnostic and box-handling changes. |
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
The malformed-line tests matched only "token", so they passed even when the error no longer named the row or reported the counts. Removing the line from both messages left all nine green, which left the fix's central guarantee — that the error identifies the annotation that failed — unprotected. Each test now matches the full message, so the quoted row and the expected/actual token counts are both pinned. The mixed valid/truncated case asserts the truncated row is the one named. With the diagnostic removed all nine fail.
docs/theme/main.html publishes this value as the page's JSON-LD dateModified, so an entry added without bumping it leaves the metadata pointing at the previous edit.
…box-extent # Conflicts: # docs/changelog.md
# Conflicts: # docs/changelog.md
…box-extent # Conflicts: # docs/changelog.md
…box-extent # Conflicts: # docs/changelog.md
|
Addressed the Copilot note on the malformed-line tests. They matched only Also merged |
…box-extent # Conflicts: # docs/changelog.md
# Conflicts: # docs/changelog.md
…-line --- Co-authored-by: claude[bot] <209825114+claude[bot]@users.noreply.github.com>
- roboflow#4: docs(changelog): drop duplicate roboflow#2663 entry left by the merge - roboflow#7: reject nan and infinite YOLO box values with one finite-value error before the extent check - roboflow#8: make YOLO line dispatch exhaustive so every parsed line appends exactly one box [resolve group] PR roboflow#2665 — items 4 7 8 --- Co-authored-by: claude[bot] <209825114+claude[bot]@users.noreply.github.com> Co-authored-by: Codex <codex@openai.com>
- roboflow#9: pin from_yolo error wrapping for an invalid label line and skipping of trailing blank lines - roboflow#10: cover the ten-token OBB line with a trailing confidence in the token-count test - roboflow#12: fold blank and whitespace-only lines into the too-short YOLO line test [resolve group] PR roboflow#2665 — items 9 10 12 --- Co-authored-by: claude[bot] <209825114+claude[bot]@users.noreply.github.com> Co-authored-by: Codex <codex@openai.com>
…ording - roboflow#13: document the invalid-line ValueError in the YOLO loader and from_yolo Raises sections - roboflow#14: docs(changelog): correct the roboflow#2665 entry's malformed-line error description [resolve group] PR roboflow#2665 — items 13 14 --- Co-authored-by: claude[bot] <209825114+claude[bot]@users.noreply.github.com> Co-authored-by: Codex <codex@openai.com>
…trings - roboflow#16: docs(changelog): move roboflow#2665 entry to top of Unreleased with wrapped link - roboflow#17: state the box-alignment invariant and test scenarios instead of past failures [resolve group] PR roboflow#2665 — items 16 17 --- Co-authored-by: claude[bot] <209825114+claude[bot]@users.noreply.github.com> Co-authored-by: Codex <codex@openai.com>
Co-authored-by: Jirka Borovec <6035284+Borda@users.noreply.github.com>
…shape (#2665) _parse_box built xyxy as x_center +/- width/2 without checking the sign, so a label such as '0 0.5 0.5 -0.4 0.4' produced [70, 30, 30, 70]: x_min past x_max. Nothing downstream expects that ordering and nothing reported it. Detections.area came back -1600, box_iou_batch scored two boxes covering the identical region at 0.0 instead of 1.0, and with_nms kept both of a pair it should have suppressed, so a corrupt label silently skewed evaluation. Reject a negative extent on load, matching how this parser already rejects a class id that is not a whole number and a non-numeric trailing token. Detections does not enforce x_min <= x_max, so object_to_yolo could write the negative extent the loader now refuses, leaving supervision unable to read a dataset it had just written. Order the corners before measuring the width and height so both ends agree; the exported box describes the same rectangle and normal boxes are byte-identical. A zero extent is degenerate but still ordered, so it keeps loading. The polygon path is unaffected: polygon_to_xyxy takes min and max. NaN and inf extents were already rejected by Detections' finite-value check. --------- Co-authored-by: jirka <6035284+borda@users.noreply.github.com> Co-authored-by: claude[bot] <209825114+claude[bot]@users.noreply.github.com> Co-authored-by: Codex <codex@openai.com>

Description
Fixes #2664. Stacked on #2663, which touches the same function; the diff below is only this change.
yolo_annotations_to_detectionsappended a class id for every line but only appended a box when the line matched a known shape, with noelse. A line of one to four tokens therefore leftclass_id_listone entry longer thanrelative_xyxy_list, and the mismatch surfaced far from its cause:0 0.5 0.5 0.2operands could not be broadcast together with shapes (0,) (4,)IndexError: list index out of rangeclass_id must be a 1D np.ndarray with shape (1,), but got shape (2,)is_obb=True, 6 / 8 / 10 tokenscannot reshape array of size 5 into shape (2)The third is what a user sees when one row of an otherwise good file is damaged, and it blames
class_idrather than the short line. None of the four name the offending line, so finding it in a large dataset means bisecting by hand.After:
This matches how the parser already rejects a class id that is not a whole number and a non-numeric trailing token.
Scope
is_obb=Trueis documented as nine-token four-corner lines only, so any other count is rejected rather than reaching the reshape.Type of change
A file that previously raised one of the errors above now raises a clearer one; nothing that previously loaded stops loading.
How has this change been tested?
Nine new tests in
tests/dataset/formats/test_yolo.py: class-only, one, two and three coordinate lines; a blank line; a truncated line among valid ones; and OBB lines that are too few, odd or too many.Following AGENTS.md §8, with a baseline captured before any change:
The failing set is identical in both runs — all eight are pre-existing
ImportError: metrics extra is requiredin this environment.uv run pre-commit runpasses on all changed files, mypy included.Docs
docs/changelog.mdentry added, linking this PR.