Client recipes
Build small clients that handle responses, errors and the appropriate execution contract.
In this topic
These recipes start with model discovery, make one explicit request, and handle failure before adding more features. Use the server address from your app. The examples assume a local server at port 1337 and a model identifier you have already verified.
A compatible SDK connects to Mellow's HTTP surface; it does not authenticate you to a cloud provider or automatically install a model. For route scope and lifecycle, read HTTP API.
Shared setup
Use two application configuration values: MELLOW_BASE_URL for the API base and MELLOW_MODEL for a discovered model. When the server requires authentication, supply MELLOW_ACCESS_KEY through a secret store. Never ship a real access key in browser JavaScript.
curl -sS http://127.0.0.1:1337/v1/models
Copy an exact returned model identifier into your configuration. Avoid examples that hardcode a model family your installation may not support.
Python with an OpenAI-compatible client
Install the openai package in your project's environment, then create one reusable client:
import os
from openai import OpenAI
client = OpenAI(
base_url=os.environ.get("MELLOW_BASE_URL", "http://127.0.0.1:1337/v1"),
api_key=os.environ.get("MELLOW_ACCESS_KEY", "local-client"),
timeout=120.0,
max_retries=0,
)
model = os.environ["MELLOW_MODEL"]
reply = client.chat.completions.create(
model=model,
messages=[{"role": "user", "content": "Explain how to organize a research project."}],
)
print(reply.choices[0].message.content or "")
local-client is only an SDK placeholder for a server configuration that permits the loopback request. It is not a Mellow credential. Replace it through the environment whenever the server requires authentication.
For streaming, request stream=True and accumulate deltas:
stream = client.chat.completions.create(
model=model,
messages=[{"role": "user", "content": "Give me three research questions."}],
stream=True,
)
try:
for event in stream:
for choice in event.choices:
if choice.delta.content:
print(choice.delta.content, end="", flush=True)
finally:
stream.close()
Close the stream when the user cancels. A UI should retain partial text while clearly distinguishing cancellation from successful completion.
Node.js without an SDK
This uses the runtime's fetch implementation and reports HTTP errors before decoding a successful result:
const base = process.env.MELLOW_BASE_URL ?? "http://127.0.0.1:1337/v1";
const model = process.env.MELLOW_MODEL;
if (!model) throw new Error("Set MELLOW_MODEL to a discovered model ID");
const controller = new AbortController();
const timer = setTimeout(() => controller.abort(), 120_000);
try {
const headers = { "Content-Type": "application/json" };
if (process.env.MELLOW_ACCESS_KEY) {
headers.Authorization = `Bearer ${process.env.MELLOW_ACCESS_KEY}`;
}
const response = await fetch(`${base}/chat/completions`, {
method: "POST",
headers,
signal: controller.signal,
body: JSON.stringify({
model,
messages: [{ role: "user", content: "Summarize the purpose of project notes." }],
stream: false,
}),
});
if (!response.ok) throw new Error(`Mellow returned HTTP ${response.status}`);
const result = await response.json();
console.log(result.choices?.[0]?.message?.content ?? "");
} finally {
clearTimeout(timer);
}
This is a server-side recipe. In a public web application, put the Mellow connection and credential behind your own authenticated backend. Browser CORS and permission boundaries still apply even when a URL is reachable.
Swift with URLSession
A native client can use the same JSON contract. Keep networking outside the main UI update path and update observable state only after decoding the result.
import Foundation
struct ChatReply: Decodable {
struct Choice: Decodable {
struct Message: Decodable { let content: String? }
let message: Message
}
let choices: [Choice]
}
func askMellow(baseURL: URL, model: String, prompt: String, key: String?) async throws -> String {
var request = URLRequest(url: baseURL.appendingPathComponent("chat/completions"))
request.httpMethod = "POST"
request.timeoutInterval = 120
request.setValue("application/json", forHTTPHeaderField: "Content-Type")
if let key {
request.setValue("Bearer \(key)", forHTTPHeaderField: "Authorization")
}
request.httpBody = try JSONSerialization.data(withJSONObject: [
"model": model,
"messages": [["role": "user", "content": prompt]],
"stream": false
])
let (data, response) = try await URLSession.shared.data(for: request)
guard let response = response as? HTTPURLResponse,
(200..<300).contains(response.statusCode) else {
throw URLError(.badServerResponse)
}
return try JSONDecoder().decode(ChatReply.self, from: data)
.choices.first?.message.content ?? ""
}
Pass a base such as http://127.0.0.1:1337/v1. This small decoder handles text-only success. Extend it to preserve tool calls, finish reasons, usage, and typed error bodies before using it as a general client.
Keep conversation history explicitly
For compatible chat requests, retain the ordered user and assistant messages your application wants to send. Append the returned assistant message, then append the next user message. Set a stable session_id only when you need Mellow's supported session bookkeeping; it is not a replacement for sending the required context.
For tool calling, preserve the assistant message containing the tool calls and attach each result to its matching call identifier. Do not insert a guessed answer when a tool failed. Return the actual structured failure so the next model step can recover honestly.
Delegate a task instead of implementing the loop
If you want Mellow to execute an existing agent's tools, discover the agent and use /agents/{id}/dispatch. Store the returned task ID and poll location. Poll with bounded backoff; render clarification as a waiting state; expose cancellation; and stop polling at a terminal state.
Do not automatically repeat dispatch after a timeout without checking whether the original task was created. A retry could launch duplicate work. Networked clients must satisfy the secure-channel and scope requirements of the selected host.
Add memory attribution deliberately
Use X-Mellow-Agent-Id with a real discovered identifier when attributing supported requests. Import historical material only through the explicit ingestion flow and only when the user intends to retain it. Keep memory configuration and extraction readiness visible rather than assuming every successful chat request becomes a pinned memory.
Test before broadening the integration
Verify one text request, an unavailable model, an authentication failure, cancellation, a second turn, and a real tool-result continuation. Then add the chosen media or agent-task contract. Keep server transport, model behavior, permissions, and task completion as separate checks in your application diagnostics.
Continue exploring · Build with MellowConfiguration reference →Understand settings ownership, local paths and supported runtime overrides.