Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
1 change: 1 addition & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -127,6 +127,7 @@ You write a `SKILL.md`. Forge compiles it into a secure, runnable agent with egr
| [Configuration](docs/reference/forge-yaml-schema.md) | `forge.yaml` schema and environment variables |
| [Settings](docs/reference/settings.md) | Layered developer settings (user + managed/MDM) — enabled channels, model default + gateway, builtin tools |
| [Dashboard](docs/reference/web-dashboard.md) | Web UI features and architecture |
| [Multimodal I/O](docs/reference/multimodal-io.md) | Sending images/PDFs to an agent and receiving media back over A2A |
| [Deployment](docs/deployment/kubernetes.md) | Container packaging, Kubernetes, air-gap |
| [Scheduler — Kubernetes](docs/deployment/scheduler-kubernetes.md) | Hybrid file/CronJob scheduler backend, RBAC, token plumbing |
| [Hooks](docs/core-concepts/hooks.md) | Agent loop hook system |
Expand Down
108 changes: 108 additions & 0 deletions docs/reference/multimodal-io.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,108 @@
---
title: "Multimodal I/O over A2A"
description: "Sending images and PDFs to a forge agent, and receiving media back, over the A2A protocol."
order: 9
---

# Multimodal I/O over A2A

Forge agents accept **images** and **PDF documents** as input and can return media as output — over the **existing A2A message schema**. No envelope change was required: A2A already modeled media as a `file` part; forge now honors those parts instead of dropping them (#255).

> **Backward compatible.** Text-only messages are unchanged and serialize byte-for-byte as before. Media support is purely additive.

## The A2A shape (unchanged)

A message is a list of typed parts. The media type is `file`:

```jsonc
// a2a.Part
{ "kind": "file", "file": { "name": "photo.png", "mimeType": "image/png", "bytes": "<base64>" } }
```

`FileContent` fields: `name` (optional), `mimeType` (required for media), `bytes` (the raw file **standard-base64-encoded**), or `uri` (by reference).

> **Inline `bytes` only, today.** Forge feeds media to the model from inline `bytes`. A `file` part with only a `uri` (no bytes) is **not fetched** — it's rejected. URI-fetch (with egress control) is a tracked follow-up. Send `bytes`.

## Input — send an image or PDF

### JSON-RPC (`POST /`)

```json
{
"jsonrpc": "2.0", "id": 1, "method": "tasks/send",
"params": {
"id": "t-1",
"message": {
"role": "user",
"parts": [
{ "kind": "text", "text": "What's in this image?" },
{ "kind": "file", "file": { "name": "photo.png", "mimeType": "image/png", "bytes": "iVBORw0KGgo..." } }
]
}
}
}
```

### REST (`POST /tasks/send`)

Same `message`, wrapped in the REST envelope:

```json
{ "task": { "id": "t-1", "message": { "role": "user", "parts": [ /* …same parts… */ ] } } }
```

A PDF is identical with `"mimeType": "application/pdf"`.

### What the model must support

Media the resolved model can't consume is **rejected loudly** (a 4xx / JSON-RPC error naming the part + reason) — never silently dropped and answered anyway.

| Media | MIME | Capable models |
|-------|------|----------------|
| Image | `image/png`, `image/jpeg`, `image/gif`, `image/webp` | vision models — OpenAI `gpt-4o`/`gpt-4.1`/`gpt-5`/`o1`/`o3`/`o4`, Anthropic Claude 3+, Gemini 1.5/2 |
| PDF | `application/pdf` | Anthropic Sonnet 3.5+, Opus 4+, Haiku 4.5+, Fable 5 |

### Limits (enforced at ingest)

| Bound | Limit | Reject reason |
|-------|-------|---------------|
| Per-image bytes | 5 MiB | `image_limit_exceeded` |
| Image dimensions | 100 000 px/side, 50 MP total | `image_limit_exceeded` |
| Images per message | 20 | `too_many_image_parts` |
| Per-PDF bytes | 32 MiB (+`%PDF-` sniff) | `document_limit_exceeded` |
| PDFs per message | 5 | `too_many_document_parts` |
| Request body | 32 MiB (both transports) | HTTP 413 |
| Concurrent media requests | 4 | 429 / unavailable (shed) |

A rejection emits an [`input_media_rejected`](../security/audit-logging.md) audit event. Note media **bytes aren't text-scannable**, so guardrail/intent scanning applies to the text/data parts only.

## Output — receive media back

Also the existing A2A shape: the response `message.parts` can carry `file` parts alongside text.

```json
{
"role": "agent",
"parts": [
{ "kind": "text", "text": "Here's the chart:" },
{ "kind": "file", "file": { "mimeType": "image/png", "bytes": "iVBORw0KGgo..." } }
]
}
```

File parts come from two sources:

- **Tools** that produce files (`file_create`, `browser_screenshot`) — always on.
- **Model-generated images** — opt-in via `models.default.image_generation: true` on an `openai-responses` model (sends the `image_generation` built-in tool; see [forge.yaml schema](forge-yaml-schema.md)).

## Persistence

Inbound uploads and model-generated output are written under the agent's files dir — `.forge/files/inbound/` and `.forge/files/generated/` respectively — content-addressed. Session history stores only the on-disk path (never base64), and media is reloaded per turn, so multi-turn conversations keep their media without bloating the session file. See [Runtime Engine → Image and document input](../core-concepts/runtime-engine.md#image-and-document-input-multimodal).

## In the web dashboard

The [chat UI](web-dashboard.md) has a **📎 attach** button: select one or more images/PDFs (validated against the accepted types and size caps above), send them with your message, and see images rendered inline and documents as download links — in both your message and the agent's reply.

## Follow-ups (not yet supported)

URI-fetch input (egress-gated) · OpenAI Responses `input_file` documents · Gemini/Bedrock media · text-extraction fallback for non-native document models · remote/distributed session media replay · streaming (partial-image) generated output.
2 changes: 2 additions & 0 deletions docs/reference/web-dashboard.md
Original file line number Diff line number Diff line change
Expand Up @@ -55,6 +55,8 @@ Click any running agent to open a chat interface that streams responses via the
| Markdown rendering | Code blocks, tables, lists rendered inline |
| Session history | Browse and resume previous conversations |
| Tool call visibility | See which tools the agent invokes during execution |
| Attachments (📎) | Attach images (PNG/JPEG/GIF/WebP) and PDFs to a message — validated against the accepted types and size caps before upload. Requires a model that supports the modality. See [Multimodal I/O](multimodal-io.md). |
| Media replies | Images the agent returns render inline; documents appear as download links. |

## Create Agent Wizard

Expand Down
35 changes: 29 additions & 6 deletions forge-ui/chat.go
Original file line number Diff line number Diff line change
Expand Up @@ -30,8 +30,8 @@ func (s *UIServer) handleChat(w http.ResponseWriter, r *http.Request) {
writeError(w, http.StatusBadRequest, "invalid request body")
return
}
if req.Message == "" {
writeError(w, http.StatusBadRequest, "message is required")
if req.Message == "" && len(req.Attachments) == 0 {
writeError(w, http.StatusBadRequest, "message or attachment is required")
return
}

Expand All @@ -53,6 +53,8 @@ func (s *UIServer) handleChat(w http.ResponseWriter, r *http.Request) {
sessionID = fmt.Sprintf("%s-%d", agentID, time.Now().UnixNano())
}

parts := buildChatParts(req.Message, req.Attachments)

// Build A2A JSON-RPC request for tasks/sendSubscribe.
rpcBody, err := json.Marshal(map[string]any{
"jsonrpc": "2.0",
Expand All @@ -61,10 +63,8 @@ func (s *UIServer) handleChat(w http.ResponseWriter, r *http.Request) {
"params": map[string]any{
"id": sessionID,
"message": map[string]any{
"role": "user",
"parts": []map[string]any{
{"kind": "text", "text": req.Message},
},
"role": "user",
"parts": parts,
},
},
})
Expand Down Expand Up @@ -152,6 +152,29 @@ func (s *UIServer) handleChat(w http.ResponseWriter, r *http.Request) {
flusher.Flush()
}

// buildChatParts projects the chat message text plus any attachments into A2A
// message parts: the text (if non-empty) as a `text` part, then a `file` part
// per attachment. Each attachment's Data is already standard-base64 — exactly
// what the agent's FileContent.Bytes ([]byte) decodes from — so it passes
// straight through onto the wire (#255).
func buildChatParts(message string, attachments []ChatAttachment) []map[string]any {
parts := make([]map[string]any, 0, 1+len(attachments))
if message != "" {
parts = append(parts, map[string]any{"kind": "text", "text": message})
}
for _, att := range attachments {
parts = append(parts, map[string]any{
"kind": "file",
"file": map[string]any{
"name": att.Name,
"mimeType": att.MimeType,
"bytes": att.Data,
},
})
}
return parts
}

// handleListSessions returns stored chat sessions for an agent.
func (s *UIServer) handleListSessions(w http.ResponseWriter, r *http.Request) {
agentID := r.PathValue("id")
Expand Down
83 changes: 83 additions & 0 deletions forge-ui/chat_media_test.go
Original file line number Diff line number Diff line change
@@ -0,0 +1,83 @@
package forgeui

import (
"net/http"
"net/http/httptest"
"strings"
"testing"
)

// TestBuildChatParts verifies the chat → A2A part projection: text becomes a
// text part, each attachment a file part with its base64 bytes passed through
// unchanged (#255).
func TestBuildChatParts(t *testing.T) {
t.Run("text only", func(t *testing.T) {
parts := buildChatParts("hello", nil)
if len(parts) != 1 || parts[0]["kind"] != "text" || parts[0]["text"] != "hello" {
t.Fatalf("parts = %+v, want one text part", parts)
}
})

t.Run("text + image attachment", func(t *testing.T) {
parts := buildChatParts("look", []ChatAttachment{
{Name: "p.png", MimeType: "image/png", Data: "aW1nLWJhc2U2NA=="},
})
if len(parts) != 2 {
t.Fatalf("parts = %+v, want [text, file]", parts)
}
if parts[1]["kind"] != "file" {
t.Fatalf("second part kind = %v, want file", parts[1]["kind"])
}
file, ok := parts[1]["file"].(map[string]any)
if !ok {
t.Fatalf("file payload malformed: %+v", parts[1]["file"])
}
if file["mimeType"] != "image/png" || file["name"] != "p.png" {
t.Errorf("file meta = %+v", file)
}
if file["bytes"] != "aW1nLWJhc2U2NA==" {
t.Errorf("base64 bytes must pass through unchanged; got %v", file["bytes"])
}
})

t.Run("attachment only (no text) omits the text part", func(t *testing.T) {
parts := buildChatParts("", []ChatAttachment{{MimeType: "application/pdf", Data: "JVBERi0="}})
if len(parts) != 1 || parts[0]["kind"] != "file" {
t.Fatalf("parts = %+v, want a single file part", parts)
}
})
}

// TestHandleChat_RequiresMessageOrAttachment: the relaxed guard — a message
// with neither text nor attachments is a 400; an attachment-only message passes
// validation (and then fails later only because no agent is running).
func TestHandleChat_RequiresMessageOrAttachment(t *testing.T) {
s, _ := newTestServer(t)

t.Run("empty message and no attachments → 400", func(t *testing.T) {
req := httptest.NewRequest(http.MethodPost, "/api/agents/test-agent/chat", strings.NewReader(`{"message":""}`))
req.SetPathValue("id", "test-agent")
rec := httptest.NewRecorder()
s.handleChat(rec, req)
if rec.Code != http.StatusBadRequest {
t.Fatalf("status = %d, want 400", rec.Code)
}
if !strings.Contains(rec.Body.String(), "message or attachment") {
t.Errorf("body = %s", rec.Body.String())
}
})

t.Run("attachment-only passes validation (fails later on no running agent)", func(t *testing.T) {
body := `{"message":"","attachments":[{"mimeType":"image/png","data":"eA=="}]}`
req := httptest.NewRequest(http.MethodPost, "/api/agents/test-agent/chat", strings.NewReader(body))
req.SetPathValue("id", "test-agent")
rec := httptest.NewRecorder()
s.handleChat(rec, req)
// Past the message/attachment guard; the next gate (agent not running)
// fires instead — i.e. the attachment-only request was NOT rejected for
// lacking a message.
if strings.Contains(rec.Body.String(), "message or attachment") {
t.Errorf("attachment-only must pass the message guard; got %s", rec.Body.String())
}
})
}
Loading
Loading