Summary
With the in-process transport (in-process / bundled-in-process features), FfiHost::start_blocking boxes a CallbackState and passes the raw pointer as user_data to copilot_runtime_connection_open. If the call returns 0, the SDK reclaims the pointer right away with Box::from_raw, drops it, and then calls host_shutdown.
Elsewhere, the SDK treats user_data as reclaimable only once the runtime has signalled quiescence. FfiShared::close / release_callback_state free the state only after copilot_runtime_connection_close returns true, and retry on a background thread until it does (the fix from #2610 / #2622 in v1.0.14). The failed-open path was not covered by that fix.
The C ABI as documented in this repo does not justify the immediate free. The prototype comment in go/internal/ffihost/ffihost.go and ADR-007 (java/docs/adr/adr-007-native-bundling-strategy.md) say only that connection_open "registers the on_outbound callback" and "Returns a connection handle (0 = failure)". Neither promises that a failed open did not retain user_data or will not invoke the callback with it. A failed open also returns no connection id, so the host has nothing to close and no quiescence signal to wait for.
If the runtime has already set up outbound delivery before the open fails, a callback can arrive after the SDK freed the state. on_outbound would then dereference a freed CallbackState and send on a dropped tx.
The Go SDK is not exposed to this: it passes an opaque token as user_data and removes it from its lookup map on failure, so a late callback is a harmless miss. The Rust SDK passes a raw heap pointer, so it is exposed.
Affected versions
- SDK rust v1.0.14 and v1.0.15. The failed-open arm in
rust/src/ffi.rs is byte-identical in both.
- Observed against CLI runtime libraries 1.0.84-5 through 1.0.89.
Code location
rust/src/ffi.rs, impl FfiHost { fn start_blocking }, at rust/v1.0.15:
let state_ptr = Box::into_raw(Box::new(CallbackState { tx, closing: AtomicBool::new(false) }));
let connection_id = unsafe { (self.connection_open)(server_id, on_outbound, state_ptr as *mut c_void, /* … */) };
if connection_id == 0 {
drop(unsafe { Box::from_raw(state_ptr) });
unsafe { (self.host_shutdown)(server_id) };
return Err(Error::with_message(ErrorKind::InvalidConfig, "copilot_runtime_connection_open failed"));
}
Compare FfiShared::close / release_callback_state in the same file. They free only after connection_close returns true.
Minimal reproduction
This is a timing-dependent memory-safety issue, so the reliable repro is a shim library plus a sanitizer:
- Build a host with
features = ["in-process"] against a shim library. The shim's copilot_runtime_connection_open spawns a thread that calls on_outbound(user_data, bytes, len) after a short delay, then returns 0. The other exports forward to a real runtime library.
- Start the client and observe
Err("copilot_runtime_connection_open failed").
- Under ASan or Miri-style tooling, the delayed callback reads the freed
CallbackState and sends on a dropped tx.
Expected vs actual
- Expected: after a failed open, the SDK does not free
user_data unless the C ABI guarantees it was not retained.
- Actual: it is freed synchronously, while the ABI as documented leaves open whether the runtime still holds it and may invoke the callback with it.
Proposed fix
Minimal fix, which we carry as a local patch: on the connection_id == 0 arm, do not reclaim state_ptr. Deliberately leak the CallbackState, which is one small struct plus an unbounded-channel sender per failed open. In practice only host boot retries produce failed opens. Keep host_shutdown and the error return unchanged. Keep release_callback_state as the single place that calls Box::from_raw.
Alternatives:
- Pass an opaque token as
user_data and look it up in a registry, as the Go SDK does. A late callback then finds no entry.
- Document in the shared C ABI that a
0 return from connection_open never retains or invokes user_data, with the runtime guaranteeing it. The current free would then be sound as written.
Patch diff summary
rust/src/ffi.rs: 1 hunk. It removes 1 line (drop(unsafe { Box::from_raw(state_ptr) });) and adds a comment explaining the intentional leak. No API change. Afterwards Box::from_raw appears exactly once in the file, inside release_callback_state.
Summary
With the in-process transport (
in-process/bundled-in-processfeatures),FfiHost::start_blockingboxes aCallbackStateand passes the raw pointer asuser_datatocopilot_runtime_connection_open. If the call returns0, the SDK reclaims the pointer right away withBox::from_raw, drops it, and then callshost_shutdown.Elsewhere, the SDK treats
user_dataas reclaimable only once the runtime has signalled quiescence.FfiShared::close/release_callback_statefree the state only aftercopilot_runtime_connection_closereturnstrue, and retry on a background thread until it does (the fix from #2610 / #2622 in v1.0.14). The failed-open path was not covered by that fix.The C ABI as documented in this repo does not justify the immediate free. The prototype comment in
go/internal/ffihost/ffihost.goand ADR-007 (java/docs/adr/adr-007-native-bundling-strategy.md) say only thatconnection_open"registers theon_outboundcallback" and "Returns a connection handle (0 = failure)". Neither promises that a failed open did not retainuser_dataor will not invoke the callback with it. A failed open also returns no connection id, so the host has nothing to close and no quiescence signal to wait for.If the runtime has already set up outbound delivery before the open fails, a callback can arrive after the SDK freed the state.
on_outboundwould then dereference a freedCallbackStateand send on a droppedtx.The Go SDK is not exposed to this: it passes an opaque token as
user_dataand removes it from its lookup map on failure, so a late callback is a harmless miss. The Rust SDK passes a raw heap pointer, so it is exposed.Affected versions
rust/src/ffi.rsis byte-identical in both.Code location
rust/src/ffi.rs,impl FfiHost { fn start_blocking }, atrust/v1.0.15:Compare
FfiShared::close/release_callback_statein the same file. They free only afterconnection_closereturnstrue.Minimal reproduction
This is a timing-dependent memory-safety issue, so the reliable repro is a shim library plus a sanitizer:
features = ["in-process"]against a shim library. The shim'scopilot_runtime_connection_openspawns a thread that callson_outbound(user_data, bytes, len)after a short delay, then returns 0. The other exports forward to a real runtime library.Err("copilot_runtime_connection_open failed").CallbackStateand sends on a droppedtx.Expected vs actual
user_dataunless the C ABI guarantees it was not retained.Proposed fix
Minimal fix, which we carry as a local patch: on the
connection_id == 0arm, do not reclaimstate_ptr. Deliberately leak theCallbackState, which is one small struct plus an unbounded-channel sender per failed open. In practice only host boot retries produce failed opens. Keephost_shutdownand the error return unchanged. Keeprelease_callback_stateas the single place that callsBox::from_raw.Alternatives:
user_dataand look it up in a registry, as the Go SDK does. A late callback then finds no entry.0return fromconnection_opennever retains or invokesuser_data, with the runtime guaranteeing it. The current free would then be sound as written.Patch diff summary
rust/src/ffi.rs: 1 hunk. It removes 1 line (drop(unsafe { Box::from_raw(state_ptr) });) and adds a comment explaining the intentional leak. No API change. AfterwardsBox::from_rawappears exactly once in the file, insiderelease_callback_state.