Skip to content

bug(ci): avoid exporting traces to an absent Kubernetes collector #3767

Description

@elezar

User Story

As an OpenShell maintainer, I want Kubernetes CI to enable OTLP export only when a usable trace collector is present, so that CI logs are actionable and the trace instrumentation is validated instead of silently dropping every exported batch.

Problem Statement

The Kubernetes E2E harness always layers deploy/helm/openshell/ci/values-skaffold.yaml into its Helm deployment. That file configures the gateway OTLP endpoint as http://openshell-collector.observability.svc.cluster.local:4317, which is appropriate for the local k3s development workflow that installs the collector.

The GitHub Actions kind workflow does not install that collector. The gateway and its in-process Kubernetes driver therefore attempt to export trace batches to a DNS name that does not exist in the CI cluster and repeatedly log BatchSpanProcessor.ExportError messages. Trace export failures are nonfatal by design, so the E2E tests continue while every batch is dropped.

Impact / Why This Matters

The repeated error messages add substantial noise to Kubernetes E2E diagnostics and can be mistaken for the cause of unrelated failures. CI currently exercises the exporter failure path accidentally, but does not verify that gateway or compute-driver traces can actually reach a collector.

The current workaround is to ignore these errors or manually override the OTLP endpoint. That is insufficient because it normalizes error-level noise in CI and leaves regressions in trace emission, propagation, resource attributes, and collector compatibility undetected.

Acceptance Criteria

  • Kubernetes CI lanes that do not validate tracing render no OTLP endpoint and do not emit repeated collector DNS/export errors.
  • Any CI lane with OTLP export enabled provisions or targets a reachable collector.
  • At least one Kubernetes CI path verifies that the collector receives traces for a representative gateway request and Kubernetes compute-driver operation.
  • Collector-backed validation fails clearly when expected traces are missing rather than silently accepting dropped batches.
  • Local k3s/Skaffold development retains its collector-backed tracing workflow.
  • The relevant E2E and observability documentation describes which workflows enable or validate OTLP export.

Reproduction Steps

  1. Run a Kubernetes E2E task through the GitHub Actions kind workflow, such as e2e:kubernetes:workspace-managed.
  2. Observe that the Helm deployment includes the OTLP endpoint from deploy/helm/openshell/ci/values-skaffold.yaml.
  3. Confirm that the CI cluster has no openshell-collector service in the observability namespace.
  4. Inspect the gateway pod logs and observe export failures at each batch interval.

Environment

Logs

[pod/openshell-0/openshell-gateway] ERROR opentelemetry_sdk: name="BatchSpanProcessor.ExportError" error="Operation failed: TonicTracesClient export failed with gRPC code: Unavailable: transport error: dns error: dns error: failed to lookup address information: Name or service not known"

Related Context

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    area:buildRelated to CI/CD and buildsarea:clusterRelated to running OpenShell on k3s/dockertopic:observabilityLogging, metrics, and observability work

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions