Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
55 changes: 39 additions & 16 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -39,7 +39,7 @@ Creating a message queue publisher:
```csharp
var options = new QueueOptions(
queueName: "my-queue",
bytesCapacity: 1024 * 1024);
capacity: 1024 * 1024);

using var publisher = factory.CreatePublisher(options);
publisher.TryEnqueue(message);
Expand All @@ -50,7 +50,7 @@ Creating a message queue subscriber:
```csharp
options = new QueueOptions(
queueName: "my-queue",
bytesCapacity: 1024 * 1024);
capacity: 1024 * 1024);

using var subscriber = factory.CreateSubscriber(options);
subscriber.TryDequeue(messageBuffer, cancellationToken, out var message);
Expand All @@ -71,7 +71,7 @@ Creating a message queue publisher using an instance of `IQueueFactory` retrieve
```csharp
var options = new QueueOptions(
queueName: "my-queue",
bytesCapacity: 1024 * 1024);
capacity: 1024 * 1024);

using var publisher = factory.CreatePublisher(options);
publisher.TryEnqueue(message);
Expand All @@ -82,7 +82,7 @@ Creating a message queue subscriber using an instance of `IQueueFactory` retriev
```csharp
var options = new QueueOptions(
queueName: "my-queue",
bytesCapacity: 1024 * 1024);
capacity: 1024 * 1024);

using var subscriber = factory.CreateSubscriber(options);
subscriber.TryDequeue(messageBuffer, cancellationToken, out var message);
Expand Down Expand Up @@ -112,7 +112,7 @@ Please note that you can start multiple publishers and subscribers sending and r

A lot has gone into optimizing the implementation of this library. For instance, it is mostly heap-memory allocation free, reducing the need for garbage collection induced pauses.

**Summary**: A full enqueue followed by a dequeue takes `~250 ns` on Linux, `~650 ns` on macOS, and `~300 ns` on Windows.
**Latest native macOS measurement**: a three-byte enqueue/dequeue round trip with a reused buffer averaged **210.0 ns** on an Apple M5 Max. Only the macOS results below were refreshed on September 12, 2026; the Windows and Linux sections retain their historical measurements.

**Details**: To benchmark the performance and memory usage, we use [BenchmarkDotNet][BenchmarkOrg] and perform the following runs:

Expand All @@ -124,7 +124,13 @@ A lot has gone into optimizing the implementation of this library. For instance,
You can replicate the results by running the following command:

```sh
dotnet run Interprocess.Benchmark.csproj -c Release
dotnet run --project src/Interprocess.Benchmark -c Release -- --filter '*QueueBenchmark*'
```

To compare throughput with one subscriber versus four concurrent subscribers:

```sh
dotnet run --project src/Interprocess.Benchmark -c Release -- --filter '*SubscriberBenchmark*' --iterationCount 8
```

---
Expand Down Expand Up @@ -152,20 +158,37 @@ Results:

### On macOS

Host:
Measured **September 12, 2026**, running directly on the Mac:

```text
BenchmarkDotNet v0.14.0, macOS Sequoia 15.2 (24C101) [Darwin 24.2.0]
Apple M3 Max, 1 CPU, 16 logical and 16 physical cores
.NET SDK 9.0.101
[Host] : .NET 9.0.0 (9.0.24.52809), Arm64 RyuJIT AdvSIMD
.NET 9.0 : .NET 9.0.0 (9.0.24.52809), Arm64 RyuJIT AdvSIMD
BenchmarkDotNet v0.15.8, macOS Tahoe 26.6.2 (25G83) [Darwin 25.6.0]
Apple M5 Max, 1 CPU, 18 logical and 18 physical cores
.NET SDK 10.0.401
.NET runtime 10.0.12, Arm64 RyuJIT
Release build; 3 warm-up iterations; 8 measured iterations; 1 launch
```

| Method | Mean (ns) | Error (ns) | StdDev | Gen0 | Allocated |
|-------------------------------------------------- |----------:|-----------:|-------:|---------:|----------:|
| 'Message enqueue and dequeue' | `249.2` | `0.74` | `0.62` | `-` | `-` |
| 'Message enqueue and dequeue - no message buffer' | `252.1` | `4.10` | `3.83` | `0.0038` | `32 B` |
All seven cases completed. Times are means in nanoseconds, normalized per operation. For enqueue/dequeue rows, an operation is one complete round trip. Concurrent-delivery rows report amortized time per delivered message.

| Workload | Mean (ns) | StdDev (ns) | Allocated per operation |
| --- | ---: | ---: | ---: |
| Enqueue, 3 bytes | 182.3 | 5.01 | 0 B |
| Enqueue + dequeue, 3 bytes, reused buffer | 210.0 | 0.46 | 0 B |
| Enqueue + dequeue, 3 bytes, new result array | 214.9 | 0.81 | 32 B |
| Enqueue + dequeue, 50 bytes, reused buffer | 214.6 | 1.33 | 0 B |
| Enqueue + dequeue, 50 bytes, ring-wrap workload | 223.8 | 1.08 | 0 B |
| Concurrent delivery, 8 bytes, 1 subscriber | 246.5 | 2.20 | Not measured |
| Concurrent delivery, 8 bytes, 4 subscribers | 344.7 | 1.81 | Not measured |

The enqueue case batches 320,000 messages and drains the queue outside the timed body. The ring-wrap case uses a 120-byte queue so padded 64-byte records repeatedly cross the end of the buffer; two round trips per invocation are normalized to one. Concurrent delivery uses one publisher and dedicated subscriber threads to transfer batches of 65,536 messages, including worker startup and completion in the timing.

These are in-process microbenchmarks, not end-to-end latency between separate applications. The concurrent cases measure throughput under contention, not individual message latency; their allocations were not measured. See the [complete native Mac reports and methodology](docs/benchmarks/2026-09-12/README.md) for source revision, errors, and reproduction details.

Run all cases from the repository root:

```sh
dotnet run --project src/Interprocess.Benchmark -c Release -- --filter '*' --warmupCount 3 --iterationCount 8 --artifacts BenchmarkDotNet.Artifacts
```

---

Expand Down
31 changes: 31 additions & 0 deletions docs/benchmarks/2026-09-12/README.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,31 @@
# Native macOS benchmark measurements — September 12, 2026

All seven benchmark cases were run directly on an Apple M5 Max using benchmark and library source at [`c7d0755`](https://github.com/cloudtoid/interprocess/commit/c7d07559a48045c82b63940da9f7910b888d9e82). The later README update does not change the measured code.

[Complete BenchmarkDotNet reports](macos-local.md)

## Environment

- macOS Tahoe 26.6.2 (25G83), Darwin 25.6.0.
- Apple M5 Max, arm64, 18 logical and physical cores reported.
- BenchmarkDotNet 0.15.8; .NET SDK 10.0.401; .NET runtime 10.0.12; Release configuration.
- Three warm-up iterations and eight measured iterations per case, one launch, with BenchmarkDotNet's usual pilot, overhead, and outlier handling.
- These are in-process microbenchmarks. They do not measure end-to-end latency between separate applications, idle CPU usage, or latency percentiles.

## Workloads

- **Enqueue:** 320,000 three-byte messages per invocation, normalized to one enqueue. Queue capacity is 5,120,000 bytes. Draining and validation happen outside the timed body. Each enqueue checks that it succeeded.
- **Three-byte round trips:** enqueue followed by dequeue on the calling thread, with either a reused buffer or a newly allocated result array; queue capacity 128 bytes.
- **50-byte round trips:** the same calling-thread pattern with a reused buffer and a 128-byte queue.
- **Ring-wrap workload:** 50-byte messages in a 120-byte queue, so padded 64-byte records repeatedly cross the end of the ring. Each invocation performs two enqueue/dequeue pairs; reported time is normalized to one pair.
- **Concurrent delivery:** one publisher sends 65,536 eight-byte messages to one or four dedicated subscriber threads sharing a 65,536-byte queue. Each subscriber receives an equal share. Time is normalized per delivered message and includes per-batch worker startup and completion.

Memory diagnostics reported 0 B/op for enqueue and reused-buffer round trips, and 32 B/op for the three-byte result-array case. Allocation diagnostics were not enabled for concurrent delivery, which creates workers and tasks per batch.

## Reproduce

From the repository root:

```sh
dotnet run --project src/Interprocess.Benchmark -c Release -- --filter '*' --warmupCount 3 --iterationCount 8 --artifacts BenchmarkDotNet.Artifacts
```
80 changes: 80 additions & 0 deletions docs/benchmarks/2026-09-12/macos-local.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,80 @@
# Local macOS — Apple M5 Max

Measured September 12, 2026, using source commit [`c7d0755`](https://github.com/cloudtoid/interprocess/commit/c7d07559a48045c82b63940da9f7910b888d9e82).

See [methodology and reproduction commands](README.md). The tables below are BenchmarkDotNet exports.

## EnqueueBenchmark

```

BenchmarkDotNet v0.15.8, macOS Tahoe 26.6.2 (25G83) [Darwin 25.6.0]
Apple M5 Max, 1 CPU, 18 logical and 18 physical cores
.NET SDK 10.0.401
[Host] : .NET 10.0.12 (10.0.12, 10.0.1226.42308), Arm64 RyuJIT armv8.0-a
.NET 10.0 : .NET 10.0.12 (10.0.12, 10.0.1226.42308), Arm64 RyuJIT armv8.0-a

Job=.NET 10.0 Runtime=.NET 10.0 InvocationCount=1
IterationCount=8 UnrollFactor=1 WarmupCount=3

```
| Method | Mean | Error | StdDev | Allocated |
|------------------ |---------:|--------:|--------:|----------:|
| 'Message enqueue' | 182.3 ns | 9.57 ns | 5.01 ns | - |

## QueueBenchmark

```

BenchmarkDotNet v0.15.8, macOS Tahoe 26.6.2 (25G83) [Darwin 25.6.0]
Apple M5 Max, 1 CPU, 18 logical and 18 physical cores
.NET SDK 10.0.401
[Host] : .NET 10.0.12 (10.0.12, 10.0.1226.42308), Arm64 RyuJIT armv8.0-a
.NET 10.0 : .NET 10.0.12 (10.0.12, 10.0.1226.42308), Arm64 RyuJIT armv8.0-a

Job=.NET 10.0 Runtime=.NET 10.0 IterationCount=8
WarmupCount=3

```
| Method | Mean | Error | StdDev | Gen0 | Allocated |
|-------------------------------------------------- |---------:|--------:|--------:|-------:|----------:|
| 'Message enqueue and dequeue - no message buffer' | 214.9 ns | 1.54 ns | 0.81 ns | 0.0038 | 32 B |
| 'Message enqueue and dequeue' | 210.0 ns | 1.03 ns | 0.46 ns | - | - |

## QueueExtendedBenchmark

```

BenchmarkDotNet v0.15.8, macOS Tahoe 26.6.2 (25G83) [Darwin 25.6.0]
Apple M5 Max, 1 CPU, 18 logical and 18 physical cores
.NET SDK 10.0.401
[Host] : .NET 10.0.12 (10.0.12, 10.0.1226.42308), Arm64 RyuJIT armv8.0-a
.NET 10.0 : .NET 10.0.12 (10.0.12, 10.0.1226.42308), Arm64 RyuJIT armv8.0-a

Job=.NET 10.0 Runtime=.NET 10.0 IterationCount=8
WarmupCount=3

```
| Method | Mean | Error | StdDev | Allocated |
|--------------------------------------------------- |---------:|--------:|--------:|----------:|
| 'Message enqueue and dequeue - long message' | 214.6 ns | 2.55 ns | 1.33 ns | - |
| 'Message enqueue and dequeue - ring-wrap workload' | 223.8 ns | 2.43 ns | 1.08 ns | - |

## SubscriberBenchmark

```

BenchmarkDotNet v0.15.8, macOS Tahoe 26.6.2 (25G83) [Darwin 25.6.0]
Apple M5 Max, 1 CPU, 18 logical and 18 physical cores
.NET SDK 10.0.401
[Host] : .NET 10.0.12 (10.0.12, 10.0.1226.42308), Arm64 RyuJIT armv8.0-a
ShortRun : .NET 10.0.12 (10.0.12, 10.0.1226.42308), Arm64 RyuJIT armv8.0-a

Job=ShortRun IterationCount=8 LaunchCount=1
WarmupCount=3

```
| Method | SubscriberCount | Mean | Error | StdDev |
|------------------------- |---------------- |---------:|--------:|--------:|
| **ReceiveConcurrentlyAsync** | **1** | **246.5 ns** | **4.21 ns** | **2.20 ns** |
| **ReceiveConcurrentlyAsync** | **4** | **344.7 ns** | **3.45 ns** | **1.81 ns** |
2 changes: 1 addition & 1 deletion src/Interprocess.Benchmark/Program.cs
Original file line number Diff line number Diff line change
Expand Up @@ -4,5 +4,5 @@ namespace Cloudtoid.Interprocess.Benchmark;

public sealed class Program
{
public static void Main() => _ = BenchmarkRunner.Run(typeof(Program).Assembly);
public static void Main(string[] args) => BenchmarkSwitcher.FromAssembly(typeof(Program).Assembly).Run(args);
}
23 changes: 15 additions & 8 deletions src/Interprocess.Benchmark/Queue/EnqueueBenchmark.cs
Original file line number Diff line number Diff line change
Expand Up @@ -3,11 +3,12 @@

namespace Cloudtoid.Interprocess.Benchmark;

[SimpleJob(RuntimeMoniker.Net90)]
[SimpleJob(RuntimeMoniker.Net10_0)]
[MemoryDiagnoser]
[MarkdownExporterAttribute.GitHub]
public class EnqueueBenchmark
{
private const int MessageCount = 320000;
private static readonly byte[] Message = [100, 110, 120];
private static readonly Memory<byte> MessageBuffer = new byte[Message.Length];
#pragma warning disable CS8618
Expand All @@ -19,8 +20,8 @@ public class EnqueueBenchmark
public void Setup()
{
var queueFactory = new QueueFactory();
publisher = queueFactory.CreatePublisher(new QueueOptions("qn", Path.GetTempPath(), 5120000));
subscriber = queueFactory.CreateSubscriber(new QueueOptions("qn", Path.GetTempPath(), 5120000));
publisher = queueFactory.CreatePublisher(new QueueOptions("qn", Path.GetTempPath(), MessageCount * 16));
subscriber = queueFactory.CreateSubscriber(new QueueOptions("qn", Path.GetTempPath(), MessageCount * 16));
}

[GlobalCleanup]
Expand All @@ -33,15 +34,21 @@ public void Cleanup()
[IterationCleanup]
public void DrainQueue()
{
for (int i = 8; i < 320000; i++)
subscriber.Dequeue(MessageBuffer, default);
for (var i = 0; i < MessageCount; i++)
{
if (!subscriber.TryDequeue(MessageBuffer, default, out _))
throw new InvalidOperationException("The benchmark did not enqueue the expected number of messages.");
}
}

// Expecting that there are NO managed heap allocations.
[Benchmark(Description = "Message enqueue (320,000 times)")]
[Benchmark(Description = "Message enqueue", OperationsPerInvoke = MessageCount)]
public void Enqueue()
{
for (int i = 8; i < 320000; i++)
publisher.TryEnqueue(Message);
for (var i = 0; i < MessageCount; i++)
{
if (!publisher.TryEnqueue(Message))
throw new InvalidOperationException("The benchmark queue is full.");
}
}
}
2 changes: 1 addition & 1 deletion src/Interprocess.Benchmark/Queue/QueueBenchmark.cs
Original file line number Diff line number Diff line change
Expand Up @@ -3,7 +3,7 @@

namespace Cloudtoid.Interprocess.Benchmark;

[SimpleJob(RuntimeMoniker.Net90)]
[SimpleJob(RuntimeMoniker.Net10_0)]
[MemoryDiagnoser]
[MarkdownExporterAttribute.GitHub]
public class QueueBenchmark
Expand Down
26 changes: 17 additions & 9 deletions src/Interprocess.Benchmark/Queue/QueueExtendedBenchmark.cs
Original file line number Diff line number Diff line change
Expand Up @@ -3,7 +3,8 @@

namespace Cloudtoid.Interprocess.Benchmark;

[SimpleJob(RuntimeMoniker.Net90)]
[SimpleJob(RuntimeMoniker.Net10_0)]
[MemoryDiagnoser]
[MarkdownExporterAttribute.GitHub]
public class QueueExtendedBenchmark
{
Expand All @@ -14,13 +15,11 @@ public class QueueExtendedBenchmark
private ISubscriber subscriber;
#pragma warning restore CS8618

[GlobalSetup]
public void Setup()
{
var queueFactory = new QueueFactory();
publisher = queueFactory.CreatePublisher(new QueueOptions("qn", Path.GetTempPath(), 128));
subscriber = queueFactory.CreateSubscriber(new QueueOptions("qn", Path.GetTempPath(), 128));
}
[GlobalSetup(Target = nameof(EnqueueDequeue_LongMessage))]
public void Setup() => SetupQueue(128);

[GlobalSetup(Target = nameof(EnqueueDequeue_WrappedMessages))]
public void SetupWrapped() => SetupQueue(120);

[GlobalCleanup]
public void Cleanup()
Expand All @@ -38,7 +37,9 @@ public ReadOnlyMemory<byte> EnqueueDequeue_LongMessage()
return subscriber.Dequeue(MessageBuffer, default);
}

[Benchmark(Description = "Message enqueue and dequeue - wrapped message in circular buffer")]
// A padded message occupies 64 bytes. A 120-byte ring makes message bodies cross
// the end of the buffer; a 128-byte ring only cycles between aligned slots.
[Benchmark(Description = "Message enqueue and dequeue - ring-wrap workload", OperationsPerInvoke = 2)]
public ReadOnlyMemory<byte> EnqueueDequeue_WrappedMessages()
{
if (!publisher.TryEnqueue(Message))
Expand All @@ -51,4 +52,11 @@ public ReadOnlyMemory<byte> EnqueueDequeue_WrappedMessages()

return subscriber.Dequeue(MessageBuffer, default);
}

private void SetupQueue(long capacity)
{
var queueFactory = new QueueFactory();
publisher = queueFactory.CreatePublisher(new QueueOptions("qn", Path.GetTempPath(), capacity));
subscriber = queueFactory.CreateSubscriber(new QueueOptions("qn", Path.GetTempPath(), capacity));
}
}
Loading