Skip to content
Build with Mellow

HTTP API

Discover models, call inference endpoints or choose the agent execution lifecycle.

In this topic
Choose the execution contract
  1. 01Your client
  2. 02Model endpoint or agent endpoint
  3. 03Configured model and access rules
  4. 04Response to your client

Raw model endpoints return tool calls for the client to execute. Authorized agent execution follows its own tool policies.

Mellow's local HTTP server lets an application request model output, call exposed tools, and work with Mellow agents. Choose the contract that matches your application: a compatible chat request supplies model input; an agent run delegates execution to the configured agent; a detached task returns an identifier that you monitor separately.

The routes below are implemented in the current server source. Their presence does not mean every backing model, media service, or workspace is configured on a particular Mac.

Establish the server address

The usual local base is http://127.0.0.1:1337. Confirm the address displayed by your running app, or use mellow status and mellow doctor. A configured port or MELLOW_PORT in the CLI environment can change the port used by your tools.

curl -sS http://127.0.0.1:1337/health
curl -sS http://127.0.0.1:1337/v1/models

A successful health response verifies reachability. Listing models identifies requestable model identifiers. Send a small request to an actual returned identifier before adding streaming, media, or tools.

Choose a compatible protocol

ProtocolRequest routeTypical client
Chat CompletionsPOST /v1/chat/completionsOpenAI-compatible SDK
Text completionsPOST /v1/completionsCompletion or fill-in-middle client
ResponsesPOST /v1/responsesResponses-compatible client
MessagesPOST /v1/messagesAnthropic-compatible SDK
Ollama chatPOST /api/chatOllama-compatible client
Ollama generationPOST /api/generateOllama generation client

Compatibility applies to the implemented request/response surface. Do not assume every feature of a hosted vendor API is available. Start with supported text input and expand one feature at a time. For Messages, use the registered /v1/messages route; do not invent an additional provider prefix.

Make one chat request

curl -sS http://127.0.0.1:1337/v1/chat/completions \
  -H 'Content-Type: application/json' \
  -d '{
    "model": "MODEL_ID",
    "messages": [
      {"role": "system", "content": "Explain concepts in clear, short paragraphs."},
      {"role": "user", "content": "What is a local agent workspace?"}
    ],
    "stream": false
  }'

Replace MODEL_ID with an identifier from discovery. This loopback example assumes the app permits the request without an access key; add authentication when required by your server configuration.

FieldPurpose
modelExact model identifier
messagesOrdered conversation messages
streamRequest streaming instead of a single JSON response
max_tokensExplicit output-token limit
temperature, top_pOptional sampling overrides
toolsFunction schemas offered by the client
tool_choiceTool selection instruction supported by the protocol
session_idConversation/session bookkeeping identifier

Omit sampler overrides when you intend to use the effective model and app configuration. Do not fill in arbitrary values from an unrelated model example. A stable session identifier does not automatically preserve the client's message history or force KV-cache reuse.

For a non-streamed Chat Completions response, inspect choices[].message, finish_reason, and available usage. A length-limited response is different from a naturally completed one. Handle an empty or missing content field when the message instead contains tool calls.

Read a stream correctly

Chat Completions uses server-sent events. Accumulate content deltas and tool-call deltas in order. A tool's JSON arguments may span several events; parse them only after the call is complete. Finish the response on the protocol's completion event rather than on the first text fragment.

data: {"choices":[{"index":0,"delta":{"content":"A local"},"finish_reason":null}]}

data: {"choices":[{"index":0,"delta":{"content":" workspace…"},"finish_reason":null}]}

data: [DONE]

This is a schematic excerpt, not a complete server response. Preserve protocol-specific event handling when using Responses or Messages instead of assuming every stream has the same shape. Close the request on cancellation and verify that the server settles its active work.

Decide who executes tools

A client supplying function definitions through a compatible chat API generally owns execution of the returned calls. Validate arguments, obtain any required user approval, execute the function, and append the assistant call and matching tool result before requesting continuation.

For Mellow to own the autonomous loop, use the selected agent's run/dispatch surface. Mellow then resolves that agent's permitted tools, policy, iteration budget, and approvals. Tool availability is scoped; it is not inherited from a client having discovered a tool name elsewhere.

Run an existing agent

RequestLifecycle
GET /agentsDiscover custom agents and their metadata
POST /agents/{id}/runExecute an agent loop within the request lifecycle
POST /agents/{id}/dispatchStart a detached task and return task metadata
GET /tasks/{task_id}Inspect a detached task
DELETE /tasks/{task_id}Request cancellation
POST /tasks/{task_id}/clarifySupply an answer to a pending clarification

Use an identifier returned by agent discovery. A detached request accepts a prompt and optional title:

{
  "prompt": "Read the approved project notes and summarize the open decisions.",
  "title": "Open decisions"
}

An accepted dispatch returns an identifier and poll location. Follow the returned poll URL and inspect status until a terminal result or an explicit clarification state. Acceptance is not task completion. Respect concurrency-limit responses and avoid retrying a dispatch blindly if the first outcome is unknown.

A clarification request supplies {"response":"your answer"}. Remote run and dispatch paths require their applicable secure-channel and scope checks. A plain bearer key should not be assumed to unlock owner-level host execution.

Discover and call externally exposed tools

RequestPurpose
GET /mcp/healthProbe MCP HTTP availability
GET /mcp/toolsList enabled tools allowed for external callers
POST /mcp/callInvoke an exposed tool

The HTTP invocation body names the tool and supplies an argument object:

{
  "name": "TOOL_NAME_FROM_DISCOVERY",
  "arguments": {}
}

Replace both the name and arguments using the discovered schema. The server rejects app-only tools before execution even if a caller supplies their name directly. A tool_not_exposable response is an exposure-policy decision, not a prompt to retry under another name.

Use the discovered schema for arguments. The external list can be smaller than the app's internal registry because enabled state and exposure rules apply. The number of tools is not the number of cloud agents or models.

Inspect the returned tool envelope even when HTTP transport succeeded. For clients using MCP stdio, mellow mcp provides the supported bridge; see CLI.

Work with models and media

RequestPurpose
GET /v1/modelsOpenAI-style model discovery
GET /v1/tagsOllama-style model list
POST /api/showModel metadata and capabilities
POST /v1/embeddingsEmbedding request
POST /api/embedOllama-format embedding request
POST /v1/audio/transcriptionsSpeech transcription
GET /v1/images/modelsInstalled local image-model capabilities
POST /v1/images/generationsImage generation
POST /v1/images/editsImage editing with a supported model
POST /v1/images/upscaleImage upscaling with a supported model
POST /v1/images/cancelCancel an image job
POST /v1/videos/quoteRequest a cloud video quote
POST /v1/videos/generationsStart a quoted video job
GET /v1/videos/jobs/{id}Inspect video-job state
GET /v1/videos/jobs/{id}/contentRetrieve completed media

Each media route has prerequisites beyond server health: appropriate models, provider access, supported input, and, for quoted cloud jobs, the relevant service and consent. Follow Image generation and Voice for setup before writing a client.

Attribute and ingest memory

Use X-Mellow-Agent-Id to attribute supported requests to a discovered agent. Attribution does not bypass authentication. POST /memory/ingest accepts agent_id, conversation_id, and turns, where each turn contains user and assistant text. Optional session_date supplies historical timing; skip_extraction retains transcript input without distillation.

The memory pipeline can require a configured extraction model. An ingested turn and a recalled fact are separate outcomes. See Memory internals for lifecycle and diagnosis.

Administrative and identity routes

Configuration is managed through loopback-only /admin/config/export, /schema, /plan, and /apply operations under that prefix. Runtime diagnostics include /admin/cache-stats, /admin/generation-settings, and /admin/runtime-settings. Use Configuration for the supported document workflow.

Pairing and secure-channel routes include /pair/challenge, /pair, /pair/hello, /pair/code, /pair/unpair, /pair-invite, /secure/session, and /secure/call. Use the supported pairing client and protocol rather than manufacturing requests from a copied access link. Public reachability is not public authorization.

GET /credits/balance reports the configured cloud-routing service's balance when that service is available. It is not a local-model prerequisite and should not be used to diagnose a missing local agent.

Authentication and failure handling

Mellow-issued access keys use sk-mellow and carry scope and lifecycle information. Supply them using the supported bearer-authentication mechanism. Keep keys in a secret store or process environment, not source code, screenshots, or URLs.

Different routes have different requirements: local administrative trust, agent audience, owner scope, pairing state, or secure transport. A valid credential can still be inappropriate for an operation. Read the response body before deciding whether a failure is authentication, authorization, validation, capacity, or execution.

Failure classRecovery
Connection refusedConfirm app, server, address, and port
Authentication rejectedCheck key validity, audience, expiry, and revocation
Permission deniedCheck route scope and required transport
Invalid requestCorrect fields using the current schema
Capacity limitWait or reduce concurrent work; avoid duplicate dispatch
Model unavailableResolve bundle/provider readiness
Partial stream or timeoutDetermine whether work is still running before retrying

Browser clients also need an allowed origin and must not expose a reusable host credential in publicly served JavaScript. A same-origin application backend is often the appropriate place to hold that credential.

Implementation reference: Packages/MellowCore/Networking/HTTPHandler.swift. Client recipes show minimal integrations without implying that every vendor SDK operation is supported.

Continue exploring · Build with MellowCommand-line workflows →Operate the local server and tools with the commands supported by this build.