Summary
The Batch Update page can currently show a combination such as:
Similarity 100%
Minor Change
This is misleading.
After comparing current main with commit 741b0bd2ec2f75a7e84c62fbe02654ce6bc41543, the underlying similarity calculation and risk thresholds are effectively unchanged. The regression is in the new display formatting.
Current main converts the stored similarity to an integer percentage using:
Math.round(similarity * 100)
The service stores the similarity to three decimal places. Therefore values such as:
0.995 → 99.5% → displayed as 100%
0.997 → 99.7% → displayed as 100%
0.999 → 99.9% → displayed as 100%
Those values are not actually 1.0.
The old UI at 741b0bd... displayed the underlying codeSimilarity value directly, so a value such as 0.999 remained visibly distinguishable from exact equality.
This should be fixed so 100% is reserved for an actual similarity score of 1.0.
Observed behavior
A Batch Update row can display:
Similarity 100%
while the same row is classified as:
Minor Change
This creates an obvious semantic conflict for users.
“100% similarity” normally means the compared content is identical. If the UI simultaneously says there is a code change, users are left unsure which signal to trust.
Comparison with 741b0bd...
Similarity calculation
The similarity calculation itself is not the new regression.
Both the old commit and current main use the same general getSimilarityScore() implementation based on string-similarity-js.
Both also store the score after truncating it to three decimal places:
const score = await getSimilarityScore(oldCode, newCode);
return +(Math.floor(score * 1000) / 1000);
So examples such as 0.995, 0.998, and 0.999 are valid stored values.
Risk thresholds
The current risk thresholds also match the older behavior:
| Similarity |
Classification |
< 0.75 |
Major Change |
0.75 – < 0.95 |
Noticeable Change |
>= 0.95 |
Minor Change |
Therefore a stored score of 0.999 correctly falls into Minor Change.
What changed
At 741b0bd..., the tooltip displayed the stored item.codeSimilarity value directly.
So a score of:
was still visibly different from:
Current main instead formats it as:
Math.round(similarity * 100)
which turns:
That formatting change loses meaningful information near the upper boundary.
Why this matters
Similarity is being used as a review/risk signal.
Users may use it to decide whether they should:
- inspect the detailed diff;
- trust an update as very small;
- investigate an unexpected change;
- proceed with a batch update.
Showing 100% has a much stronger meaning than showing 99.9%.
It should not be produced merely as a rounding artifact.
The current presentation can make a changed script appear fully identical even though the underlying score explicitly says otherwise.
Expected behavior
The displayed percentage should preserve the distinction between:
- exact
1.0;
- near-perfect values such as
0.995–0.999.
The important invariant should be:
Display 100% only when the underlying stored similarity is exactly 1.0.
Examples:
| Stored score |
Current UI |
Expected UI |
0.950 |
95% |
95% or 95.0% |
0.994 |
99% |
99.4% |
0.995 |
100% |
99.5% |
0.998 |
100% |
99.8% |
0.999 |
100% |
99.9% |
1.000 |
100% |
100% |
Suggested formatting
Because the service already stores similarity to three decimal places, one decimal place in percentage form preserves the same useful precision:
0.999 → 99.9%
0.995 → 99.5%
0.950 → 95.0%
1.000 → 100%
The implementation can hide a trailing .0 if desired.
For example, conceptually:
function formatSimilarity(similarity: number): string {
if (similarity === 1) return "100%";
const percent = similarity * 100;
return String(percent.toFixed(1)) + "%";
}
The exact helper implementation can differ, but it should maintain the invariant that a score below 1 never rounds upward to 100%.
A dedicated formatter is preferable to embedding number formatting directly inside RiskBadge, because the rule is semantic rather than purely cosmetic.
Do not change the risk thresholds as part of this fix
The evidence does not currently indicate that the 0.75 / 0.95 classification thresholds are the regression.
Those thresholds already existed in 741b0bd....
This issue should therefore focus on the inaccurate percentage presentation rather than changing the meaning of:
- Major Change;
- Noticeable Change;
- Minor Change.
Changing those thresholds would be a separate product/design decision.
Exact 1.0 consideration
If a future reproducible case shows that two meaningfully different script bodies produce an actual stored score of exactly 1.000, that would indicate a separate similarity-algorithm issue.
That should not be conflated with the deterministic display bug described here.
This issue can be reproduced solely from the formatter:
stored similarity = 0.999
current display = 100%
No scoring-algorithm failure is required.
Testing
Please add tests around the formatting boundary.
At minimum:
A direct component/formatter test is preferable to relying only on a screenshot test.
Acceptance criteria
Design intent
The percentage and the risk badge describe the same comparison from two different angles.
They should reinforce each other rather than appear contradictory.
“Minor Change · Similarity 99.9%” is understandable.
“Minor Change · Similarity 100%” suggests that either the change label or the percentage is wrong.
Preserving the actual near-100% value restores the distinction that the older Batch Update UI already had.
Summary
The Batch Update page can currently show a combination such as:
This is misleading.
After comparing current
mainwith commit741b0bd2ec2f75a7e84c62fbe02654ce6bc41543, the underlying similarity calculation and risk thresholds are effectively unchanged. The regression is in the new display formatting.Current
mainconverts the stored similarity to an integer percentage using:The service stores the similarity to three decimal places. Therefore values such as:
0.995→99.5%→ displayed as 100%0.997→99.7%→ displayed as 100%0.999→99.9%→ displayed as 100%Those values are not actually
1.0.The old UI at
741b0bd...displayed the underlyingcodeSimilarityvalue directly, so a value such as0.999remained visibly distinguishable from exact equality.This should be fixed so 100% is reserved for an actual similarity score of 1.0.
Observed behavior
A Batch Update row can display:
while the same row is classified as:
This creates an obvious semantic conflict for users.
“100% similarity” normally means the compared content is identical. If the UI simultaneously says there is a code change, users are left unsure which signal to trust.
Comparison with
741b0bd...Similarity calculation
The similarity calculation itself is not the new regression.
Both the old commit and current
mainuse the same generalgetSimilarityScore()implementation based onstring-similarity-js.Both also store the score after truncating it to three decimal places:
So examples such as
0.995,0.998, and0.999are valid stored values.Risk thresholds
The current risk thresholds also match the older behavior:
< 0.750.75 – < 0.95>= 0.95Therefore a stored score of
0.999correctly falls into Minor Change.What changed
At
741b0bd..., the tooltip displayed the storeditem.codeSimilarityvalue directly.So a score of:
was still visibly different from:
Current
maininstead formats it as:which turns:
That formatting change loses meaningful information near the upper boundary.
Why this matters
Similarity is being used as a review/risk signal.
Users may use it to decide whether they should:
Showing 100% has a much stronger meaning than showing 99.9%.
It should not be produced merely as a rounding artifact.
The current presentation can make a changed script appear fully identical even though the underlying score explicitly says otherwise.
Expected behavior
The displayed percentage should preserve the distinction between:
1.0;0.995–0.999.The important invariant should be:
Examples:
0.9500.9940.9950.9980.9991.000Suggested formatting
Because the service already stores similarity to three decimal places, one decimal place in percentage form preserves the same useful precision:
The implementation can hide a trailing
.0if desired.For example, conceptually:
The exact helper implementation can differ, but it should maintain the invariant that a score below
1never rounds upward to100%.A dedicated formatter is preferable to embedding number formatting directly inside
RiskBadge, because the rule is semantic rather than purely cosmetic.Do not change the risk thresholds as part of this fix
The evidence does not currently indicate that the
0.75 / 0.95classification thresholds are the regression.Those thresholds already existed in
741b0bd....This issue should therefore focus on the inaccurate percentage presentation rather than changing the meaning of:
Changing those thresholds would be a separate product/design decision.
Exact
1.0considerationIf a future reproducible case shows that two meaningfully different script bodies produce an actual stored score of exactly
1.000, that would indicate a separate similarity-algorithm issue.That should not be conflated with the deterministic display bug described here.
This issue can be reproduced solely from the formatter:
No scoring-algorithm failure is required.
Testing
Please add tests around the formatting boundary.
At minimum:
0.95does not display as 100%.0.994displays below 100%.0.995displays below 100%.0.999displays below 100%.1.0displays as exactly 100%.A direct component/formatter test is preferable to relying only on a screenshot test.
Acceptance criteria
1.0can never be presented as 100%.99.x%from100%.1.0continues to display as 100%.0.995–0.999rounding boundary.Design intent
The percentage and the risk badge describe the same comparison from two different angles.
They should reinforce each other rather than appear contradictory.
Preserving the actual near-100% value restores the distinction that the older Batch Update UI already had.