Skip to content

Add KADABRA betweenness centrality approximation algorithm - #518

Open
DerSchmachtin wants to merge 1 commit into
JuliaGraphs:masterfrom
DerSchmachtin:master
Open

DerSchmachtin wants to merge 1 commit into
JuliaGraphs:masterfrom
DerSchmachtin:master

Conversation

@DerSchmachtin

@DerSchmachtin DerSchmachtin commented Jul 31, 2026 •

Copy link
Copy Markdown

Description

This Pull Request introduces the KADABRA (K-ADaptive Approximation of Betweenness centRAlity) algorithm to Graphs.jl.

Computing exact betweenness centrality can be prohibitively expensive on very large graphs. KADABRA provides a fast and highly scalable randomized approximation of the betweenness centrality, as well as an efficient method to find the top-k nodes with the highest betweenness.

The implementation includes two main exported functions:

  • kadabra_centrality: Approximates the betweenness centrality for all vertices.
  • kadabra_top_k: Returns the top k vertices with the highest betweenness centrality.

Reference:

Borassi M., Natale E. (2016) KADABRA is an ADaptive Algorithm for Betweenness via Random Approximation. In: ESA 2016. [https://dl.acm.org/doi/10.1145/3284359]

A note on the top-k allocation

In top-k mode KADABRA sizes each vertex's confidence interval from its rank gaps before sampling begins. Here the implementation follows the paper (Section 5.2) rather than the authors' C++ reference, which NetworKit's KadabraBetweenness also copies — so a reviewer comparing against either will find a deliberate difference rather than a porting bug:

  • The reference pairs the two rank gaps opposite to the way its own stopping test consults them, and leaves lambda_L(v_1) unconstrained — pinning the top vertex's lower budget, the only one its own test uses, at the minimum.
  • Its tie-collapse rule is never applied to the pair (v_k, v_{k+1}), the single pair the external-exclusion test is built around. Neither vertex's budgets are then aimed at the err fallback, so once their gap falls below 2*err they can only be separated with bounds tighter than err, and the run oversamples.

The allocation is a heuristic for where to spend the delta budget, not part of the (err, delta) guarantee, so each variant remains a valid approximation and they differ only in how many samples they need. Measured over 4 SNAP graphs × k ∈ {3, 5, 10, 100} × 3 seeds at err = 1e-4, the paper's allocation needed 0.80 ± 0.18 times as many samples as the reference's, at identical top-k accuracy — the same overlap with the exact ranking and the same Kendall τ over the top k to three decimals. Collapsing the boundary pair additionally removes the cases where a top-k query cost more than computing every centrality. See the number of samples in comparison to 'k=0' below. k = 0 is unaffected: it does not enter this code path.

topk_allocation

I am separately seeking confirmation from the KADABRA authors on whether the reference's ordering was deliberate (I'm pretty sure it was overlooked), and will report back here. In the meantime I am happy to switch to the reference's behaviour instead if you would prefer bug-for-bug compatibility with NetworKit.

Checklist

  • Added kadabra.jl to src/centrality/
  • Added corresponding test suite in test/centrality/kadabra.jl and registered it in runtests.jl
  • Code has been formatted using JuliaFormatter.jl according to BlueStyle
  • Added documentation strings and exported the functions in Graphs.jl
  • Ensured RNG scoping follows Graphs.jl conventions using standard library imports

All tests pass locally. Let me know if there are any changes or further optimizations you would like me to make!

Benchmarks & Performance

To verify the performance, I ran some benchmarks comparing this Julia implementation against the original, highly optimized C++ version from the authors.

Hardware Setup:
All benchmarks were executed on an AMD Ryzen Threadripper 3960X 24-Core Processor (3.80 GHz, 48 threads) with 125 GiB of RAM running Ubuntu 22.04 LTS.

The Julia implementation is competitive compared to the Original C++ Implementation.
Here are the results (using k=0, delta=0.1 and epsilon=0.0001) on a few test instances taken from SNAP:

image

@DerSchmachtin
DerSchmachtin force-pushed the master branch 2 times, most recently from a71b033 to 47fd306 Compare July 31, 2026 18:26
@codecov

codecov Bot commented Jul 31, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 97.62%. Comparing base (dffc7a6) to head (221c19e).

Additional details and impacted files
@@            Coverage Diff             @@
##           master     #518      +/-   ##
==========================================
+ Coverage   97.47%   97.62%   +0.15%     
==========================================
  Files         128      129       +1     
  Lines        7811     8350     +539     
==========================================
+ Hits         7614     8152     +538     
- Misses        197      198       +1     

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@DerSchmachtin
DerSchmachtin force-pushed the master branch 5 times, most recently from 1b7aaad to 6d1291d Compare August 5, 2026 07:42
@DerSchmachtin
DerSchmachtin force-pushed the master branch 5 times, most recently from 0298521 to e213fd1 Compare August 14, 2026 11:30
@LoveLow-Global

Copy link
Copy Markdown
Member

Hi there, sorry for the delay in responding. I was very busy thoughout the last couple of weeks, and haven't had the time to look into the new PRs. I will try to review this within a week. Thank you!!

@LoveLow-Global LoveLow-Global self-assigned this Aug 19, 2026
@DerSchmachtin
DerSchmachtin force-pushed the master branch 2 times, most recently from 869c0ca to f5d137c Compare September 1, 2026 22:57
@DerSchmachtin
DerSchmachtin force-pushed the master branch 2 times, most recently from da3d329 to cddb698 Compare September 16, 2026 22:47
@LoveLow-Global

Copy link
Copy Markdown
Member

Hello, sorry for the very long delay. I was quite busy with other work for a while.

I wonder if you used AI-generated code here, could you share if you used LLMs? The overall code looks good, however as far as I know, the code should be reviewed differently when there is a significant usage of LLMs in code generation, which I do not know well. Please check this comment

@DerSchmachtin

Copy link
Copy Markdown
Author

No worries, thanks for taking a look!

Yes, I used Claude (Opus 4.8 and 5, via Claude Code) for a lot of the code. It was a series of sessions over a few months while working for an uni lab project on KADABRA. The agent had direct access to the original C++ reference implementation throughout, so it was working from an established baseline (where I found a couple of places that don't match the paper, which this implementation fixes).

A main part of the uni project was a comparison to a different algorithm for betweenness centrality, so I benchmarked them extensively and in the process looked pretty heavily at the code and validated the algorithm.
The two divergence points from the C++ version (top-k allocation and burn-in normalization) were deliberate decisions driven by those empirical measurements, both of which are documented in the PR description.

I'm also writing up a report on the comparison (deadline is the 30th of September). It's not public yet, but I can send it to you if it helps with the review.

Implements the KADABRA approximation algorithm for betweenness
centrality (Borassi & Natale, 2019), including the adaptive
delta-calibration phase, and registers it in the docs, CHANGELOG
and test suite.

The top-k confidence-budget allocation follows the paper rather than the
authors' C++ reference, which transposes the two rank gaps relative to the
stopping test its own Algorithm 2 performs, leaves lambda_L(v_1)
unconstrained, and never applies the tie-collapse rule to the pair
(v_k, v_{k+1}) that the external-exclusion test is built around. The
allocation is a heuristic for where to spend the delta budget, not part of
the (err, delta) guarantee, so both readings are valid approximations and
differ only in cost. Over 4 SNAP graphs x k in {3, 5, 10, 100} x 3 seeds at
err = 1e-4, the paper's allocation needed 0.80 +- 0.18 times as many samples
as the reference's for identical top-k accuracy, and collapsing the boundary
pair additionally removes the cases where a top-k query cost more than
computing every centrality. The rationale is documented at the call sites in
`compute_bet_err!`. k = 0 does not enter this code path.

Scores and bounds are normalised by the Phase 2 pair count. This also departs
from the reference, which zeroes the burn-in counts (Probabilistic.cpp:363-368)
but then adds tau back to the divisor (`n_pairs += tau`, line 405), scaling
every score by 1 - tau/N. On six SNAP graphs at eps = 1e-4 (tau/N between 2.9%
and 7.7%) that puts the maximum error at 3.5-12.3x eps, with 40-89 vertices over
eps per graph, in the reference's own output; normalised by Phase 2 pairs it is
0.20-0.39x eps. NetworKit divides before adding tau and is unaffected.
`n_samples` still reports every pair drawn, burn-in included, and the stopping
rule compares Phase 2 counts with Phase 2 pairs either way, so sample counts and
rankings are the same under both. At the default start_factor = 100 the burn-in
is only 1% of omega on the small test graphs, so one test sets start_factor = 2,
making it half of omega, and asserts that every estimate lies within eps of the
exact value and inside its returned bounds; under the reference's normalisation
that test fails at 1.6-2.0x eps.

Sampling workers keep thread-local counts and merge them only when they test
the stopping condition, every max(1000, tau ÷ 10) pairs, instead of taking a
lock after every sample as the reference does. Two details keep that from
overshooting: a worker tests the stop flag before every sample rather than
finishing its batch, and a check divides by the pairs every worker has drawn so
far, not by completed batches only, since the counts it merges already include
unfinished ones. On four SNAP graphs at eps = 1e-4 with 8 threads over five
seeds, the port drew 1.03-1.06x the reference's samples without them and
1.004-1.012x with them, and ran 3-6% faster; Kendall tau_b stayed within 0.002.
@LoveLow-Global

Copy link
Copy Markdown
Member

Hi, thanks for the reply. I will have to ask another maintainer on how to review your pull request. Sorry for the increased delay. Also, I believe you can share the report after the 30th, if that is when it becomes public.

@LoveLow-Global

Copy link
Copy Markdown
Member

@Krastanov As far as I am aware of, when the codes commited are AI-generated, the review process has to be slightly different. Could you inform me on how to review the code in this case? I will review this PR when I have the proper guidelines to do so.
Also, as this PR ports the KADABRA algorithm with Apache-2.0 C++ reference, I wonder if it is sufficient to only mention the paper in the docstrings.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants