Skip to content

feat(providers): add Slack provider profile with gateway-managed rotating user tokens #3729

Description

@n1hility

User Story

As an operator running long-lived agents in OpenShell sandboxes that read or post to Slack on a user's behalf, I want to give a sandbox a Slack user token from a Slack app that has token rotation enabled, so that the gateway keeps the short-lived access token fresh and the sandbox never holds a long-lived Slack credential.

Problem Statement

OpenShell ships no example provider profile for Slack, and its gateway-managed oauth2_refresh_token refresh cannot report failures from Slack's token endpoint correctly.

Slack apps with token rotation enabled issue user access tokens that expire after 43200 seconds together with a single-use refresh token, redeemed at https://slack.com/api/oauth.v2.access with the app's client ID and client secret. The successful refresh response is RFC 6749 shaped and the existing engine can mint from it. The failure response is not: Slack returns HTTP 200 with {"ok": false, "error": "invalid_refresh_token"} (and similar error names) instead of a 4xx OAuth error response.

Today the gateway decides success purely on the HTTP status, so a Slack error response is treated as a malformed success body: the refresh state records oauth_invalid_success_response, recovery action investigate, and is retried every 60 seconds indefinitely. A grant whose refresh token was revoked, or whose app was uninstalled, is never parked as reauthorization_required and the operator is never told that the user must authorize again. Separately, a profile that omits max_lifetime_seconds gets the 3600 second default, which clamps Slack's 12 hour lifetime and schedules a refresh (consuming a single-use refresh token) every hour.

Impact / Why This Matters

Without this, operators who want Slack access from a sandbox must either use non-rotating Slack apps (a permanent user token that never expires, which the OpenShell credential model is designed to avoid) or run their own refresher outside the gateway and update the provider credential externally, losing the gateway's refresh ownership, stable placeholders, and status reporting. Operators who do configure oauth2_refresh_token against Slack get working rotation on the happy path but a silent, permanent retry loop on the first real failure, with no signal that reauthorization is needed; the sandbox simply loses access when the current token expires. This was observed on a live deployment: after the Slack app was uninstalled, the refresh state stayed in investigation_required until the state was reconfigured by hand.

Proposed Design

From the user's perspective:

  • providers/slack.yaml is shipped as an example profile that an operator imports with openshell provider profile import, after which openshell provider create --type slack works. The profile declares one bearer user-token credential (SLACK_USER_TOKEN) with oauth2_refresh_token refresh against https://slack.com/api/oauth.v2.access, a 43200 second lifetime cap, a refresh lead well ahead of expiry, and client_id, client_secret (secret) and refresh_token (secret) material. openshell provider refresh configure --strategy oauth2-refresh-token works with that material exactly as it does for the Google profiles.
  • When Slack's token endpoint reports an error, openshell provider refresh status shows the same outcomes operators already know: a dead or revoked refresh token parks the state as reauthorization_required with recovery action reauthorize and failure code oauth_invalid_grant, with Slack's own error name visible as the provider error subtype; a bad client ID or secret reports fix_configuration with oauth_invalid_client; Slack's transient errors are retried. Bodies that carry a token continue to mint exactly as today, and bodies with neither a token nor an error keep today's oauth_invalid_success_response outcome.
  • The documentation describes how to create the provider, configure refresh with an existing authorization's refresh token, and read the status table, including that Slack refresh tokens are single-use and should be refreshed only by the gateway.

Acceptance Criteria

  • A slack example provider profile exists under providers/ and lints clean; after importing it, provider create --type slack followed by provider refresh configure --strategy oauth2-refresh-token with client_id, client_secret, refresh_token and an expiry configures refresh without any custom profile.
  • A refresh whose response is HTTP 200 with a Slack error body such as invalid_refresh_token or token_revoked results in reauthorization_required / reauthorize / oauth_invalid_grant and is parked, not retried every minute; the Slack error name is recorded as the provider error subtype; no provider-controlled text is copied into the status message.
  • A refresh whose response is HTTP 200 with invalid_client_id or bad_client_secret results in fix_configuration / oauth_invalid_client; ratelimited and Slack's transient errors result in retry.
  • Successful refreshes against a 43200 second expires_in schedule the next refresh according to the profile's lead time rather than being clamped to 3600 seconds.
  • Existing RFC 6749 issuers are unaffected: token responses and 4xx error responses classify exactly as before, and a 2xx body with neither a token nor an error field still reports oauth_invalid_success_response.
  • Documentation for the Slack profile and for the 2xx error handling is added.

Alternatives Considered

  • Keep the provider static (non-rotating Slack apps). This is what the observed deployment did first. It works but leaves a permanent user token in the credential store, which is the situation gateway-managed refresh exists to avoid, and Slack's rotation flag is the only way to get expiring user tokens.
  • A dedicated slack refresh strategy instead of reusing oauth2_refresh_token. Slack's refresh request and successful response are standard; only the error envelope differs, so a new strategy would duplicate the engine for a classification difference and require CLI and SDK changes for no user-visible gain.
  • A profile-level opt-in for the error envelope (for example a response_format field). This keeps the shared classifier untouched for issuers that never declare it, at the cost of a proto field on the profile and refresh state. It is a reasonable follow-up if maintainers prefer an explicit switch; the proposed behavior only changes a path that today is already a failure, so an opt-in did not seem necessary for the first version.
  • Running the refresh outside the gateway and updating the provider credential externally. This gives up refresh ownership, status reporting and the stable placeholder that lets running sandboxes survive rotation.

Agent Investigation

Findings from reading the code paths involved (crate openshell-server, provider_refresh.rs):

  • request_token reads the HTTP status, routes non-2xx bodies to classify_oauth_token_error, and deserializes 2xx bodies into TokenResponse { access_token, expires_in, refresh_token }; a 2xx body without access_token fails deserialization and becomes oauth_invalid_success_response (recovery investigate, retry in 60 seconds).
  • classify_oauth_token_error recognizes the RFC 6749 error names plus Google's invalid_rapt and admin_policy_enforced; issuer-specific names are folded onto RFC failure codes with the vendor name preserved in provider_error_subtype, which is the pattern a Slack table can follow.
  • DEFAULT_MAX_LIFETIME_SECONDS is 3600 and caps expires_in when a profile leaves max_lifetime_seconds unset; next_refresh_at_ms is expires_at_ms - refresh_before_seconds, and both values are pinned into the refresh state at configure time from the profile.
  • Slack's refresh response for user tokens places access_token, refresh_token, expires_in and token_type: "user" at the top level (only the initial authorization-code exchange nests them under authed_user), so the existing TokenResponse parses it; the refresh form sends client_id, client_secret and refresh_token in the body, which Slack accepts.
  • Verified live against a Slack app with rotation enabled: connect-time and worker-driven refreshes succeed and the stable placeholder swaps under a running process; after uninstalling the app, a forced refresh with the proposed classification parked the state as reauthorization_required with subtype invalid_refresh_token.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    state:triage-neededOpened without agent diagnostics and needs triage

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions