Skip to content

Follow Ultralytics in candidate ranking, DOTAv1 scoring and YOLOv7 NMS - #25

Merged
parkjinman98 merged 3 commits into
mainfrom
jm/ops-parity-followups
Oct 1, 2026
Merged

parkjinman98 merged 3 commits into
mainfrom
jm/ops-parity-followups

Conversation

@parkjinman98

Copy link
Copy Markdown
Contributor

Summary

The original model repository is the source of truth. Aligning mblt-model-ops with it (mblt-model-ops !230) turned up three places where mblt-vision differs from it too. This PR fixes them and bumps __version__ to 0.0.7, so ops' parity pin can follow.

1. Candidate ranking (NMS and end-to-end heads)

Ultralytics 8.4.137 (validates on CUDA) Before (main) After
NMS candidate order torchvision.ops.nms sorts stably (CPU and CUDA); the argsort above max_nms is stable on CUDA argsort(descending=True): ties reorder on CPU common.descending_order (stable) on every device
End-to-end dual_topk order torch.topk (k = 300), stable on CUDA torch.topk: ties reorder on CPU descending_order
End-to-end dual_topk rows up to max_det (anchor, class) pairs capped at the number of anchors above conf_thres up to max_det pairs

The cap bites on sparse images. With 100 anchors holding three classes each above the threshold, Ultralytics keeps 300 rows and main kept 100.

AGENTS.md required the unstable sort, believing it reproduced Ultralytics. The earlier measurement behind that rule (YOLOv8m -0.00047 mAP50-95 for a stable sort, and so on) compared two tie orders, not either one against Ultralytics.

2. DOTAv1 scoring

Ultralytics Before After
Difficult objects (flag 1/2) ordinary targets: convert_dota_to_yolo_obb drops the flag ignore regions (DOTA devkit protocol) ordinary targets
Primary metric rotated mAP50 (published "mAP test 50") mAP50-95 mAP50, with mAP50-95 secondary

evaluate_dota_predictions still honours ignore regions that a caller supplies explicitly; only the loader stops producing them.

3. YOLOv7 NMS threshold

WongKinYiu/yolov7 Before After
iou_thres (YOLOv7, -x, -w6, -e6, -d6, -e6e) 0.65 (test.py --iou-thres default; README --iou 0.65) 0.6 0.65

Changes

Area File Change
Ranking helper utils/postprocess/common.py descending_order, a stable descending sort; used by rotated NMS and dual_topk
dual_topk utils/postprocess/common.py second stage keeps up to max_det pairs; both stages use descending_order
NMS candidate sorts yolo_anchor_post.py, yolo_anchorless_post.py, yolo_dflfree_post.py, yolo_nmsfree_post.py argsort(descending=True) → descending_order
DOTA loader utils/evaluation/eval_dota.py difficult objects load as targets (both label paths); flag still validated
DOTA result eval_dota.py, cli/val.py, benchmark/benchmark_vision_models.py primary mAP50; print order and score_name follow
YOLOv7 six models/YOLOv7*.yaml iou_thres: 0.65
Version mblt_vision/__init__.py 0.0.6 → 0.0.7
Docs AGENTS.md, mblt-vision skill, mblt_vision/README.md the new rules; README OBB metric, and YOLOv7 scores marked for re-measurement

Test plan

Check Result
Parity against Ultralytics 8.4.137 on CUDA (the reference), with int8-quantized, tie-heavy inputs: dual_topk (sparse multi-class, dense) and the multi-label NMS candidate path, on CPU and CUDA tensors, 6 trials each ✅ After: 24/24 identical. Before: every CPU case differed, and sparse dual_topk also differed on CUDA
tests/test_ultralytics_selection_order.py (new) ✅ Against the previous code it fails on the CPU tie order of NMS and dual_topk, and on the row cap on both devices
DOTA tests: both loader paths count difficult objects; the benchmark runner reports map50 ✅
pytest ✅ 755 passed. ⚠️ 5 failures in tests/test_onnx_yolo.py are Hub 401 downloads; they fail identically on main
pre-commit (ruff 0.6.9) and git diff --check ✅
Re-measured README scores ⏳ Not run. YOLOv7 (iou_thres), DOTA (protocol) and any MXQ result affected by tie order need re-measurement on hardware

🤖 Generated with Claude Code

parkjinman98 and others added 3 commits October 1, 2026 16:05
Two divergences from Ultralytics 8.4.137, found while aligning
mblt-model-ops with it (mblt-model-ops !230):

- Tie order. AGENTS.md required the unstable argsort(descending=True),
  on the belief that it reproduces Ultralytics. It does not:
  non_max_suppression passes up to 30000 candidates straight to
  torchvision.ops.nms, whose sort is stable on CPU and CUDA; the argsort
  above that cap and an end-to-end head's torch.topk (k = 300) are
  stable on CUDA, where Ultralytics validates. On CPU, tied scores came
  out in another order, and quantized MXQ scores tie often. Every
  candidate ranking now goes through common.descending_order, a stable
  descending sort, which gives Ultralytics' CUDA order on every device.
  The #22 measurement (YOLOv8m -0.00047 etc.) compared two tie orders,
  not either one against Ultralytics.
- dual_topk's row cap. It ranked only the anchors above the threshold
  and then capped its second stage at that anchor count, so whenever
  fewer than max_det anchors cleared the threshold every further class of
  those anchors was dropped: 100 anchors with three classes each kept 100
  rows where Detect.get_topk_index keeps 300. The second stage now keeps
  up to max_det (anchor, class) pairs.

Verified against Ultralytics 8.4.137 run on CUDA, with int8-quantized,
tie-heavy inputs, 6 trials per case: dual_topk (sparse multi-class,
dense) and the multi-label NMS candidate path are identical on CPU and
CUDA tensors (24/24). Before this change every CPU case differed, and
sparse multi-class dual_topk differed on CUDA too.

tests/test_ultralytics_selection_order.py pins both. Run against the
previous code, it fails on the CPU tie order of NMS and dual_topk and on
the row cap on both devices. AGENTS.md and the mblt-vision skill state
the new rule.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The 15 OBB models are Ultralytics checkpoints, so their own repository is
the reference, as mblt-model-ops now holds too (mblt-model-ops !230):

- Difficult objects (flag 1 or 2) load as ordinary targets. Ultralytics'
  convert_dota_to_yolo_obb keeps the eight coordinates and the class name
  and drops the flag, so its validation counts them. The loader used to
  route them into ignore regions, the DOTA devkit protocol behind the
  test server. evaluate_dota_predictions still honours ignore regions a
  caller passes explicitly; the loader produces none. The flag is still
  validated.
- Rotated mAP50 is the primary metric, the number Ultralytics publishes
  for its OBB models (mAP test 50); mAP50-95 becomes secondary.
  DOTAResult, `mblt-vision val`'s print order and the benchmark runner's
  score_name follow.

Tests: both loader paths (normalized and official val_original labels)
count a difficult object; the supplied-ignore-region test now says what
it covers; the benchmark runner reports map50. AGENTS.md, the
mblt-vision skill and mblt_vision/README.md state the rule.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
WongKinYiu/yolov7's test.py defaults to --iou-thres 0.65, and its README
reproduces every published COCO number with `--iou 0.65`; the
iou_thres=0.6 in test()'s signature is always overridden by the command
line. All six YOLOv7 models (YOLOv7, -x, -w6, -e6, -d6, -e6e) used 0.6.
mblt-model-ops makes the same change (mblt-model-ops !230). The README's
YOLOv7 scores were measured at 0.6 and are marked for re-measurement.

__version__ moves to 0.0.7: this release changes NMS ordering, end-to-end
selection and DOTAv1 scoring, so mblt-model-ops' parity pin can follow
it. AGENTS.md and the mblt-vision skill record the YOLOv7 threshold.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@chatgpt-codex-connector

Copy link
Copy Markdown

You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard.

@parkjinman98
parkjinman98 merged commit 55a993c into main Oct 1, 2026
3 checks passed
@parkjinman98
parkjinman98 deleted the jm/ops-parity-followups branch October 1, 2026 07:25
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant