Skip to content

feat: add opt-in conversation checkpoint and resume - #2341

Open
bhargavikalicheti wants to merge 1 commit into
TEN-framework:mainfrom
bhargavikalicheti:feat/conversation-checkpoint-resume
Open

bhargavikalicheti wants to merge 1 commit into
TEN-framework:mainfrom
bhargavikalicheti:feat/conversation-checkpoint-resume

Conversation

@bhargavikalicheti

Copy link
Copy Markdown

Summary

Add an initial checkpoint/resume abstraction for #2340. Applications can save explicitly serializable conversation and workflow state, then restore it into a new session after an interruption.

  • Add versioned checkpoints and participant snapshot, validation, and restore hooks.
  • Provide pluggable storage with an in-memory implementation.
  • Require application-defined resume policies and enforce expiration.
  • Validate checkpoint identity, participant sets, and state versions before restoring.
  • Include a SQLite example demonstrating recovery across separate processes.
  • Add documentation and a Windows/Linux CI workflow.

Scope and limitations

This is an opt-in proof of concept pending maintainer feedback. It does not automatically reconnect transports or checkpoint runtime resources. Applications remain responsible for authorization, consistent snapshot boundaries, and tool-operation idempotency.

Restore hooks are not transactional. If a hook fails after another participant has restored, the application must discard the partially restored session.

Validation

  • All 21 tests pass locally, including cross-process recovery.
  • Formatting and lint checks pass.
  • Package builds and installs successfully; tests also pass against the installed copy.

Refs #2340

@bhargavikalicheti

Copy link
Copy Markdown
Author

@plutoless @halajohn Hi, I’ve prepared an initial proof of concept for this issue with opt-in state hooks, versioned checkpoints, pluggable storage, expiration, and application-defined resume policies.
It includes documentation, a SQLite example demonstrating recovery across separate processes, and 21 passing tests. Restore is explicit; automatic transport reconnection and distributed failover remain outside scope.
Before expanding the implementation, I’d appreciate feedback on whether a standalone Python system package is the right home for checkpoint coordination, and whether the proposed participant hooks fit TEN’s architecture.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant