Skip to content
Build with Mellow

Inside inference

Understand model resolution, loading, streaming and request failures.

In this topic

Mellow's local runtime translates a chat or API request into a model load, a generation plan, a stream of output, and cleanup. Model selection, concurrency, and cache controls affect different stages. This chapter gives developers a way to inspect those stages without treating every model family as interchangeable.

Follow the lifecycle

  1. Resolve the requested model bundle and its metadata.
  2. Combine model configuration with explicit request and user settings.
  3. Check the requested configuration and memory constraints before unsafe allocation.
  4. Acquire the model engine and retain a lease while generation is active.
  5. Tokenize the prompt, reuse compatible cached state, and prefill remaining input.
  6. Decode text and structured model output into the caller's stream.
  7. Finish or cancel, release the lease, and apply idle residency policy.

A downloaded model, a loaded engine, and a completed answer are three different milestones. Tool calling, vision, long context, and repeated turns need their own execution checks.

Generation defaults belong to the effective plan

Explicit request or agent choices can override applicable user defaults and bundle generation configuration. Leaving an override unset allows the model's own configuration to contribute. Clients should avoid populating every sampler field with arbitrary values simply because their SDK accepts them.

Record the effective temperature, top-p, top-k, min-p, output limit, and repetition settings when comparing runs. Greedy decoding and sampling are different modes; a changed sampler can invalidate a performance or quality comparison even when the model filename is unchanged.

Reasoning and tool boundaries depend on the model bundle, tokenizer, chat template, and runtime. Do not conceal a parser defect by adding forced markers or stripping suspicious output after generation. Capture the raw event and the rendered conversation so the failing boundary can be located.

Concurrency has several owners

ComponentResponsibility
Batch engineSchedule compatible requests within a model engine
Model registryCoalesce engine creation for the same model
Model leasePrevent unloading a model while a request still uses it
Residency managerDecide when an idle model should unload
GPU coordinationCoordinate producers that share Metal resources
Plugin host limitsBound inference requests initiated by a plugin

Continuous batching concerns compatible work inside an engine. It does not make every pair of models fit in memory, and it does not grant unlimited tool or plugin concurrency. Inspect queueing and active slots when a request appears to wait.

Cache reuse is about compatible state

Prefix reuse follows the effective prompt and model state, not a magic conversation identifier. Keep session_id stable for conversation bookkeeping, but do not treat it as a command to reuse incompatible cache contents. An informational prefix hash can help detect a prompt change; sending the hash back is not cache control.

Changes to system instructions, tool schemas, media, or compacted conversation context can change reuse. Context compaction may replace older outbound turns with a summary while leaving the visible transcript intact. The next generation therefore needs a cache identity consistent with the new prompt.

Different model architectures retain different state. Full-attention KV, hybrid recurrent/SSM state, and other pooling or sliding-window structures cannot be validated with the same cache metric. Disk cache availability also depends on its directory, budget, and model support.

Use runtime evidence

The local server provides administrative diagnostics including:

GET /admin/cache-stats
GET /admin/generation-settings
GET /admin/runtime-settings

Use these from the local machine under the app's administrative access rules. Compare effective settings and cache counters before and after a real request. A counter increase is useful evidence, but does not establish a coherent answer or correct tool execution by itself.

For performance, retain first-load time, time to first token, prefill rate, decode rate, peak physical memory, and the selected bundle. Separate a cold request from a warm repeat. Use a second turn that genuinely depends on the first when assessing continuity.

Cancellation and unloading

Test cancellation during initial load, prefill, generation, and tool continuation. The UI should settle, input should unlock, and resources should stop growing. Closing a stream should not leave an unowned model load running in the background.

Idle unloading begins after active leases are released. It is separate from deleting downloaded weights or clearing persistent cache. A model disappearing from resident memory is not evidence that its files were removed.

Troubleshooting matrix

SymptomInspect
Load rejectedBundle support, effective memory limit, and requested context
Slow first answerDownload completeness, load, prefill, and cold cache
Slow later turnsChanged prefix, disabled disk tier, incompatible cache state
Queue never settlesActive requests, cancellation, lease ownership
Raw tool or reasoning markersTemplate/parser contract and full stream
Answer changes after cache restoreBaseline without reuse and architecture-specific state restoration

Support should be stated per tested model and scenario. A runtime implementation or isolated unit test is not a blanket compatibility claim.

Continue exploring · Build with MellowSandbox execution →Follow runtime preparation, environment boundaries and agent code execution.