Repository navigation
Follow Ultralytics in candidate ranking, DOTAv1 scoring and YOLOv7 NMS - #25
Merged
Merged
Conversation
Two divergences from Ultralytics 8.4.137, found while aligning mblt-model-ops with it (mblt-model-ops !230): - Tie order. AGENTS.md required the unstable argsort(descending=True), on the belief that it reproduces Ultralytics. It does not: non_max_suppression passes up to 30000 candidates straight to torchvision.ops.nms, whose sort is stable on CPU and CUDA; the argsort above that cap and an end-to-end head's torch.topk (k = 300) are stable on CUDA, where Ultralytics validates. On CPU, tied scores came out in another order, and quantized MXQ scores tie often. Every candidate ranking now goes through common.descending_order, a stable descending sort, which gives Ultralytics' CUDA order on every device. The #22 measurement (YOLOv8m -0.00047 etc.) compared two tie orders, not either one against Ultralytics. - dual_topk's row cap. It ranked only the anchors above the threshold and then capped its second stage at that anchor count, so whenever fewer than max_det anchors cleared the threshold every further class of those anchors was dropped: 100 anchors with three classes each kept 100 rows where Detect.get_topk_index keeps 300. The second stage now keeps up to max_det (anchor, class) pairs. Verified against Ultralytics 8.4.137 run on CUDA, with int8-quantized, tie-heavy inputs, 6 trials per case: dual_topk (sparse multi-class, dense) and the multi-label NMS candidate path are identical on CPU and CUDA tensors (24/24). Before this change every CPU case differed, and sparse multi-class dual_topk differed on CUDA too. tests/test_ultralytics_selection_order.py pins both. Run against the previous code, it fails on the CPU tie order of NMS and dual_topk and on the row cap on both devices. AGENTS.md and the mblt-vision skill state the new rule. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The 15 OBB models are Ultralytics checkpoints, so their own repository is the reference, as mblt-model-ops now holds too (mblt-model-ops !230): - Difficult objects (flag 1 or 2) load as ordinary targets. Ultralytics' convert_dota_to_yolo_obb keeps the eight coordinates and the class name and drops the flag, so its validation counts them. The loader used to route them into ignore regions, the DOTA devkit protocol behind the test server. evaluate_dota_predictions still honours ignore regions a caller passes explicitly; the loader produces none. The flag is still validated. - Rotated mAP50 is the primary metric, the number Ultralytics publishes for its OBB models (mAP test 50); mAP50-95 becomes secondary. DOTAResult, `mblt-vision val`'s print order and the benchmark runner's score_name follow. Tests: both loader paths (normalized and official val_original labels) count a difficult object; the supplied-ignore-region test now says what it covers; the benchmark runner reports map50. AGENTS.md, the mblt-vision skill and mblt_vision/README.md state the rule. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
WongKinYiu/yolov7's test.py defaults to --iou-thres 0.65, and its README reproduces every published COCO number with `--iou 0.65`; the iou_thres=0.6 in test()'s signature is always overridden by the command line. All six YOLOv7 models (YOLOv7, -x, -w6, -e6, -d6, -e6e) used 0.6. mblt-model-ops makes the same change (mblt-model-ops !230). The README's YOLOv7 scores were measured at 0.6 and are marked for re-measurement. __version__ moves to 0.0.7: this release changes NMS ordering, end-to-end selection and DOTAv1 scoring, so mblt-model-ops' parity pin can follow it. AGENTS.md and the mblt-vision skill record the YOLOv7 threshold. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
The original model repository is the source of truth. Aligning mblt-model-ops with it (mblt-model-ops !230) turned up three places where mblt-vision differs from it too. This PR fixes them and bumps
__version__to 0.0.7, so ops' parity pin can follow.1. Candidate ranking (NMS and end-to-end heads)
main)torchvision.ops.nmssorts stably (CPU and CUDA); theargsortabovemax_nmsis stable on CUDAargsort(descending=True): ties reorder on CPUcommon.descending_order(stable) on every devicedual_topkordertorch.topk(k = 300), stable on CUDAtorch.topk: ties reorder on CPUdescending_orderdual_topkrowsmax_det(anchor, class) pairsconf_thresmax_detpairsThe cap bites on sparse images. With 100 anchors holding three classes each above the threshold, Ultralytics keeps 300 rows and
mainkept 100.AGENTS.mdrequired the unstable sort, believing it reproduced Ultralytics. The earlier measurement behind that rule (YOLOv8m -0.00047 mAP50-95 for a stable sort, and so on) compared two tie orders, not either one against Ultralytics.2. DOTAv1 scoring
convert_dota_to_yolo_obbdrops the flagevaluate_dota_predictionsstill honours ignore regions that a caller supplies explicitly; only the loader stops producing them.3. YOLOv7 NMS threshold
iou_thres(YOLOv7, -x, -w6, -e6, -d6, -e6e)test.py --iou-thresdefault; README--iou 0.65)Changes
utils/postprocess/common.pydescending_order, a stable descending sort; used by rotated NMS anddual_topkdual_topkutils/postprocess/common.pymax_detpairs; both stages usedescending_orderyolo_anchor_post.py,yolo_anchorless_post.py,yolo_dflfree_post.py,yolo_nmsfree_post.pyargsort(descending=True)→descending_orderutils/evaluation/eval_dota.pyeval_dota.py,cli/val.py,benchmark/benchmark_vision_models.pyscore_namefollowmodels/YOLOv7*.yamliou_thres: 0.65mblt_vision/__init__.pyAGENTS.md,mblt-visionskill,mblt_vision/README.mdTest plan
dual_topk(sparse multi-class, dense) and the multi-label NMS candidate path, on CPU and CUDA tensors, 6 trials eachdual_topkalso differed on CUDAtests/test_ultralytics_selection_order.py(new)dual_topk, and on the row cap on both devicesmap50pytesttests/test_onnx_yolo.pyare Hub401downloads; they fail identically onmainpre-commit(ruff 0.6.9) andgit diff --checkiou_thres), DOTA (protocol) and any MXQ result affected by tie order need re-measurement on hardware🤖 Generated with Claude Code