LLMs
ProtoLink's LLM package starts with a simple promise: choose where a model runs, then use the same application-facing contract to talk to it. That contract can stop after one text response, stream text as it arrives, or continue into a controlled Agent loop where the model requests tools and delegates work while ProtoLink remains responsible for execution.
The useful mental model is a progression:
- Choose a backend: hosted API, model server, in-process local model, or deterministic mock.
- Create one adapter with
create_llm()or a concrete provider class. - Pick the interaction level:
chat()for one direct response,infer()for controlled multi-step work, or the lower-levelcall*()methods when implementing an adapter. - Let an Agent own runtime concerns such as tool selection, delegation, policy, cancellation, state, and events.
- Add operational features when needed: persistent history, explicit compaction, context reporting, cost estimates, and run budgets.
Choose a model backend, make direct or streamed calls, and graduate to a controlled inference loop without rewriting the rest of the application.
protolink.llmscreate_llm()LLM.chat()LLM.infer()LLM.compact_history()Start with the interaction you need
Most application code should begin with either chat() or an Agent. The other methods exist so provider adapters and advanced integrations can participate in the same runtime.
-
One complete response -
llm.chat("...")
Adds the user message and performs one synchronous provider call. -
Visible incremental text -
llm.chat("...", streaming=True)
Returns an asynchronous iterator of text chunks. -
Tools, delegation, policy, retries, budgets, and multiple steps -
Agent(..., llm=llm)
The Agent prepares the runtime and invokesLLM.infer()with the appropriate services. -
A custom controlled runtime -
await llm.infer(...)
Advanced integration path where the caller prepares the prompt, history, tools, callbacks, and runtime context that an Agent normally supplies. -
A provider adapter or diagnostic integration -
call(),call_stream(), andcall_action()
Works at the provider boundary with explicitConversationHistoryand normalized actions.
chat() and infer() are intentionally different. A chat call asks the model for text. Inference asks the model for one typed decision at a time: finish with text, call a local tool, or delegate to another Agent. ProtoLink validates that decision and performs the side effect outside the model.
Use chat() for a small standalone model interaction. Use an Agent when model output can cause work to happen. Reach for call() and call_stream() primarily when extending or diagnosing a provider adapter.
The current common contract is text-oriented: chat inputs, canonical history content, complete responses, and stream chunks are strings. Provider support does not imply that every provider-specific feature is wrapped. Use the provider SDK directly for capabilities outside this surface, such as file-upload APIs, image or audio generation, embeddings, fine-tuning, batch administration, or model management.
Choose where the model runs
ProtoLink groups model backends into hosted APIs, model servers, and in-process local runtimes. A deterministic testing adapter follows the same contract without making a model request.
-
API - calls a remote provider and normally requires an API key:
OpenAILLM: OpenAI Responses API, including native function tools and streamed tool-call events.AnthropicLLM: Anthropic Messages API, including nativetool_useblocks.GeminiLLM: Google GenAI API, including native function declarations.DeepSeekLLM: DeepSeek Chat Completions API, with optional native tools.GrokLLM: xAI Chat Completions API, with optional native tools.HuggingFaceLLM: Hugging Face Inference API for non-streaming direct calls.
-
Server - connects to a model server that you run locally or remotely:
OllamaLLM: connects to an Ollama/api/chatendpoint.LlamaCPPServerLLM: connects directly to allama-server.LMStudioLLM: connects to LM Studio's OpenAI-compatible server.VLLMLLM: connects to vLLM's OpenAI-compatible server.OpenAICompatibleLLM: connects to a service exposing/v1/chat/completionsand/v1/models.
-
Local - runs the model inside the Python process:
LlamaCPPLocalLLM: loads a local GGUF file throughllama-cpp-python.
For deterministic tests and offline examples, MockLLM can return fixed responses, response sequences, or callback-generated responses. It is the fastest way to test Agent behavior without credentials, network access, or nondeterministic model output.
You can also use a third-party LLM client directly when your application only needs that client's raw API. ProtoLink's wrappers become valuable when you want provider-neutral history, streaming, typed actions, Agent integration, compaction, events, metrics, or budgets.
How to choose
- Choose a hosted API when you want a managed model and accept provider credentials, network latency, and usage billing.
- Choose a server adapter when the model is exposed by Ollama, llama-server, LM Studio, vLLM, or another OpenAI-compatible endpoint that you control.
- Choose
LlamaCPPLocalLLMwhen the GGUF model should run inside the Python process and the process can afford model-loading and inference resources. - Choose
MockLLMfor tests, examples, and runtime development.
Native tool support is an adapter-and-model capability, not a requirement for ProtoLink inference. OpenAI, Anthropic, and Gemini use provider-native tool structures. DeepSeek and Grok enable native tools by default but can fall back to the portable JSON action protocol. Self-hosted and local adapters use that JSON path by default unless a known-compatible model/server is explicitly opted into native tool calling.
Install, configure, and make the first call
Configuration varies by backend, but the first successful model interaction follows one continuous path.
-
Install the relevant extras:
# All supported LLM backendsuv add "protolink[llms]"
If you only need a subset of providers, install their SDKs directly instead of the llms extra, which installs every supported integration. Server and local adapters may not need a hosted-provider SDK.
-
Construct the LLM through the lazy factory:
from protolink import create_llmllm = create_llm("openai",model="gpt-4o-mini",# api_key is normally read from OPENAI_API_KEY)Import a concrete adapter when provider-specific configuration makes the code clearer:
from protolink.llms.api import OpenAILLMllm = OpenAILLM(model="gpt-4o-mini")The factory lazily imports only the selected adapter and its optional SDK. Construction itself is not deferred: concrete adapters currently validate during initialization and may perform network, server-health, model-loading, or filesystem work.
Never commit API keys to version control. Read them from environment variables or a secure secrets manager.
-
Make a direct call:
response = llm.chat("Explain what this service does in two sentences.")print(response)As an alternative, use a fresh or explicitly managed history and consume the asynchronous iterator returned by streaming chat:
streaming_llm = create_llm("openai", model="gpt-4o-mini")async for chunk in streaming_llm.chat("Draft a short welcome message.",streaming=True,):print(chunk, end="", flush=True) -
Pass the same LLM to an Agent when the model should participate in a controlled runtime:
from protolink.agents import Agentfrom protolink.models import AgentCardagent_card = AgentCard(url="http://localhost:8020",name="llm_agent",description="Agent backed by an LLM",)agent = Agent(card=agent_card, transport="http", llm=llm)
For local and server-style LLMs (LlamaCPPLocalLLM, LlamaCPPServerLLM, OllamaLLM, LMStudioLLM, VLLMLLM, and OpenAICompatibleLLM), configuration additionally includes a model-file path or server URL. The individual class entries below describe the resolution order for those values.
Generation parameters are deliberately provider-specific. model_params is forwarded to the selected SDK or server; ProtoLink does not translate names such as max_tokens, max_output_tokens, or provider-specific thinking controls into one synthetic schema.
chat() appends the user message before calling the provider, but it does not append the returned assistant text. That small behavior keeps the base method predictable across streaming and non-streaming adapters. Manage ConversationHistory explicitly for a standalone multi-turn script, or let an Agent own task-local and persistent conversation history.
From text generation to controlled inference
Direct generation ends when the provider returns text. Controlled inference can continue because the model is treated as a planner rather than an executor.
Each inference step follows the same story:
- ProtoLink adds the user query and prepares the current system prompt, history, tools, and discovered Agents.
- The provider adapter obtains one decision, either through native function/tool calling or the portable JSON action protocol.
- ProtoLink validates the decision as
FinalAction,ToolCallAction, orAgentCallAction. - A final action returns user-facing text. A tool or Agent action passes through authorization and is executed by the runtime.
- The result becomes a new observation in history, and the model receives another step only when more work is necessary.
answer = await agent.invoke(
"Summarize the available information.",
session_id="customer-42",
)
print(answer)
The Agent rebuilds the system prompt for the current tool set, action mode, discovered Agents, flow position, and application instructions before it invokes LLM.infer(). It also supplies the policy authorizer, cancellation token, run context, budget policy, event callback, and isolated history.
infer() directlyDirect inference is available for custom runtimes, but passing a tools dictionary does not itself rebuild the LLM's system prompt. A direct caller must call build_system_prompt() with the matching tool and Agent descriptions and must provide any callbacks, authorization, cancellation, and context it needs. For normal tool use, prefer Agent.add_tool() or @agent.tool followed by await agent.invoke(...).
Streaming means two different things
chat(..., streaming=True)returns an async iterator because the chunks are the result.infer(streaming=True)still returns one finalPart. Intermediate model text is emitted asllm_chunkevents through the inference event callback while the runtime waits for a complete, validated action.
This distinction lets an Agent show live progress without dispatching a partial or malformed tool request.
Tools and delegated Agents stay behind a runtime boundary
Provider-native adapters expose tool schemas using the provider's function-calling format. Other adapters describe the same choices in the system prompt and parse one JSON action. Both routes converge on the same typed action models before anything executes.
The runtime then applies the protections that a raw model call does not provide: unknown-tool checks, policy and approval decisions, cancellation, duplicate-action detection, parse-error feedback, transient provider retries, per-run budgets, and a ten-step inference limit.
The detailed controlled inference and tool-use chapter below explains action acquisition, recovery, prompts, and delegation. The LLM.infer() reference documents every integration hook.
Conversation history inside and outside an Agent
An LLM instance still exposes llm.history for direct usage and backward-compatible introspection. When the same LLM is plugged into an Agent, ProtoLink binds a task-local ConversationHistory around each run so concurrent tasks do not interleave messages on one shared mutable history object.
For stateless agents, each task receives a fresh history seeded by the compiled system prompt. After normal completion, llm.history points at a copy of that task history for debugging and simple scripts. A failed turn normally stays isolated; if an action_result receipt proves that a side effect completed, ProtoLink retains the history containing that observation so a retry does not behave as though the action never ran.
For persistent conversation state, enable state=["conversation"]. The Agent loads the requested session_id, serializes concurrent tasks for that same session with an async lock, saves normally completed history back to state, and exposes a copy as llm.history. Failed history is saved only when a new action_result receipt proves that the turn completed a tool or delegation side effect.
from protolink import Agent, AgentCard, RunContext, Task, create_llm
agent = Agent(
AgentCard(
name="assistant",
description="Assistant",
url="runtime://assistant",
),
llm=create_llm("mock", default_response="ok"),
state=["conversation"],
)
task = Task.create_infer(prompt="remember this")
RunContext(session_id="customer-42").attach_to_task(task)
await agent.execute_task(task)
Direct llm.infer(...) calls are unchanged: they use the LLM's default history unless you explicitly call llm.use_history(history).
History compaction
Every LLM wrapper owns a modular HistoryCompactor at llm.compactor. Its compact() method mutates the current ConversationHistory in place and returns a HistoryCompactionResult with before/after message and estimated-token counts. LLM.compact_history() remains as a convenient facade, so direct usage stays concise.
# Fastest: keep the system prompt and 19 newest messages.
report = llm.compact_history("recent", max_messages=20)
# Local and budget-aware: keep a recent suffix near 8,000 estimated tokens.
report = llm.compact_history(
"tokens",
max_tokens=8_000,
preserve_recent=6,
)
# Highest context fidelity: summarize old turns, preserve the newest 8 verbatim.
report = llm.compact_history(
"summary",
preserve_recent=8,
summary_max_tokens=600,
)
print(report.to_dict())
# Equivalent component-oriented API:
report = llm.compactor.compact("tokens", max_tokens=8_000)
Three strategies cover different cost and fidelity needs:
| Strategy | Model calls | Behavior | Best for |
|---|---|---|---|
recent | 0 | Keeps the leading system prompt and newest max_messages messages. | A simple, fast sliding window. |
tokens | 0 | Keeps the newest chronological suffix near max_tokens, using tiktoken when installed or the built-in estimate otherwise. | Deterministic context-budget control. |
summary | 1 | Replaces older turns with a model-generated system summary and keeps preserve_recent turns verbatim. | Retaining decisions and constraints from long sessions. |
The tokens limit is deliberately soft when the leading system prompt plus protected recent messages already exceed the budget: ProtoLink preserves those messages instead of silently removing the active request. The summary strategy makes its model call with a temporary history. The live history is changed only after a non-empty summary is returned, so a provider failure leaves it untouched.
Agent-requested compaction
Agent-requested compaction is a control-plane request, not a model tool and not a task part. Call the Agent method directly or use the client request spec. The compaction capability is never appended to the model prompt and is never exposed through provider-native or JSON tool calling, which keeps the prompt smaller and friendlier to very small models.
from protolink import HistoryCompactionRequest
report = await agent.compact_history(
HistoryCompactionRequest(
strategy="summary",
preserve_recent=8,
summary_max_tokens=600,
session_id="customer-42",
)
)
For remote agents, use the client spec-backed convenience method:
from protolink.client import AgentClient
client = AgentClient("runtime", url="runtime://client")
report = await client.compact_history(
"runtime://agent",
strategy="summary",
preserve_recent=8,
summary_max_tokens=600,
session_id="customer-42",
)
The remote path is POST /llm/history/compact, represented by AgentClient.COMPACT_HISTORY_REQUEST and an EndpointSpec registered by AgentServer. When state=["conversation"] is enabled and session_id is supplied, the Agent loads the session history before compaction and saves the compacted history afterward. The runtime action is still evaluated through the policy boundary with the llm.history.compact capability, so applications can allow, deny, or require approval for context loss.
ProtoLink does not compact history automatically based on an arbitrary context threshold. Applications can call llm.compact_history() for local use, agent.compact_history() inside an Agent process, or AgentClient.compact_history() over a transport. Natural-language requests such as “please compact your context” are application intent; convert them into a control-plane request before calling the Agent if you want deterministic behavior.
How the LLM package is organized
The public facade is protolink.llms.base.LLM. Internally, the base class owns provider-neutral orchestration: history binding, metrics, budgets, retries, tool execution, Agent delegation, and final response handling. The strict action parser lives in protolink.llms.parsing, where raw model text becomes one validated LLMAction and narrow fallback shorthands are repaired only when the target tool or Agent is unambiguous.
Provider adapters keep request and stream handling in their own modules, then return typed results to the shared inference loop:
protolink.llms.apicontains hosted-provider adapters.protolink.llms.servercontains HTTP model-server adapters.protolink.llms.localcontains the in-process llama.cpp adapter.protolink.llms.mock_clientprovides deterministic testing behavior.
The rest of the package is partitioned by responsibility:
factory.pylazily resolves provider names and constructs adapters.history.pydefines canonical messages and provider-neutral conversation history.actions.py,parsing.py, andtool_calling.pydefine typed runtime actions and translate provider-native function calls into them.context.pybuilds a context manifest for each inference step.compaction.pyowns explicit history reduction and summary compaction.metrics.pynormalizes usage, estimates missing token counts, and calculates application-supplied costs.prompts/keeps JSON-action and provider-native prompt families separate so the model never receives conflicting tool instructions.serialization.pyprovides JSON-safe conversion for history and runtime payloads._deps.pyloads optional provider SDKs only when their adapter is selected.
This separation keeps the everyday API small while allowing native tool-calling providers and JSON-fallback models to participate in the same controlled runtime.
Observe and constrain model work
Profiles, context manifests, call metrics, and run budgets matter when a working model integration becomes an operated system. They are intentionally optional and do not need to be configured before the first call.
Model metadata
LLMModelProfile describes the deployment information ProtoLink cannot safely hardcode: context-window size, application-supplied input and output prices, tokenizer metadata, and descriptive capability flags. A profile does not change the provider request or enable a feature in the selected model.
from protolink import LLMModelProfile, create_llm
llm = create_llm(
"openai-compatible",
model="my-model",
metrics_profile=LLMModelProfile(
context_window=128_000,
input_cost_per_million=1.0, # example value; use current provider pricing
output_cost_per_million=5.0, # example value; use current provider pricing
supports_tools=True,
supports_streaming=True,
supports_json_schema=True,
tokenizer="cl100k_base",
),
)
You can also configure metrics after construction:
llm.configure_metrics(
context_window=128_000,
input_cost_per_million=1.25,
output_cost_per_million=10.0,
)
Provider-reported token usage is used when the SDK response includes it. Otherwise ProtoLink estimates token counts locally. If tiktoken is installed through protolink[metrics], ProtoLink resolves an encoder from the active model name and falls back to cl100k_base; without it, ProtoLink uses a lightweight character heuristic. The profile's tokenizer field is currently descriptive metadata and does not select the estimator. Prices, model limits, and capabilities change over time, so LLMModelProfile is application-owned metadata rather than a hardcoded billing catalog.
Context and call events
LLM wrappers can emit pre-call context manifests plus per-call latency, token usage, context-window pressure, and estimated cost through the existing infer() event stream and telemetry hooks. Observing these events does not change the request payload sent to the provider.
Before each model call, ProtoLink emits a provider-neutral context_prepared event:
{
"type": "context_prepared",
"step": 1,
"manifest": {
"run_id": "run_123",
"agent_name": "researcher",
"system_tokens": 900,
"tool_prompt_tokens": 300,
"history_tokens": 2200,
"user_tokens": 120,
"total_estimated_tokens": 3520,
"context_window": 128000,
},
}
When an event_callback or telemetry backend is attached, each model call inside the inference loop can also emit:
{
"type": "llm_call_metrics",
"step": 1,
"provider": "openai-compatible",
"model": "my-model",
"latency_ms": 842.37,
"usage": {
"input_tokens": 1200,
"output_tokens": 180,
"estimated": False,
},
"context": {
"used_tokens": 1200,
"window_tokens": 128000,
"used_percent": 0.938,
},
"cost": {
"input_cost": 0.0012,
"output_cost": 0.0009,
"total_cost": 0.0021,
},
}
This is especially useful for CLIs, dashboards, and budget-aware agents that want to show context pressure or session cost while a multi-step tool loop is running.
Run-budget enforcement
Model profiles are observational metadata. Run budgets are the separate enforcement mechanism.
If a RunContext carries a RunBudget, LLM.infer() enforces it through the default BudgetEnforcer. When an Agent executes a task, one task-local enforcer is shared by every infer part and explicit tool-call part, so counters do not reset between parts. A nested task gets its own scope, and direct LLM.infer() calls still create an independent enforcer unless the advanced caller supplies budget_enforcer= explicitly.
Pre-call limits such as max_llm_calls and max_input_tokens are checked before every physical provider attempt, including transient retries. max_tool_calls applies before explicit and model-selected tools execute; max_output_tokens is checked after provider usage or local estimates are available. Provider runtime is checked again after each request, including a final response. A tool or delegated call can already have committed a side effect when it returns, so ProtoLink records and injects that result before cancellation or runtime enforcement stops the next step. Warnings appear as budget_warning events and hard denials appear as budget_exceeded events.
LLM API reference
The rest of this page is the detailed contract. It starts with construction and the base methods, then follows controlled inference, prompts, concrete providers, related objects, examples, and failure handling.
Provider switching in action
The same application code works across providers. Keep provider choice in configuration, construct exactly one adapter, and leave the calling code unchanged. chat() is the high-level convenience method for direct text generation; internally it selects call() or call_stream().
from protolink import create_llm
# Choose one deployment in application configuration.
provider = "ollama"
provider_options = {
"openai": {
"model": "gpt-4o-mini",
},
"anthropic": {
"model": "claude-sonnet-4-20250514",
},
"ollama": {
"model": "qwen3",
"base_url": "http://localhost:11434",
},
"lmstudio": {
"model": "local-model",
"base_url": "http://localhost:1234/v1",
},
"vllm": {
"model": "Qwen/Qwen3-8B",
"base_url": "http://localhost:8000/v1",
},
}
llm = create_llm(provider, **provider_options[provider])
# The calling code stays the same.
response = llm.chat("Hello! How are you?")
print(response)
LLM- abstract base class with the common runtime behavior.APILLM- base for API-hosted adapters.ServerLLM- base for HTTP model servers.LocalLLM- base for in-process local runtimes.- Concrete implementations -
OpenAILLM,AnthropicLLM,GeminiLLM,DeepSeekLLM,GrokLLM,HuggingFaceLLM,OllamaLLM,LlamaCPPServerLLM,LMStudioLLM,VLLMLLM,OpenAICompatibleLLM,LlamaCPPLocalLLM, andMockLLM.
Each callable below keeps the scikit-learn-style layout: exact signature, explanation, separately labeled parameters and defaults, return values, raised errors, notes, and focused examples.
create_llm
create_llm(
provider: str | LLMProvider,
**kwargs,
) -> LLMCreate an LLM adapter without importing the selected provider until it is needed.
Parameters
providerstr | LLMProviderrequiredProvider selector. Supported string values are
anthropic,deepseek,gemini,grok,huggingface,llama.cpp-local,llama.cpp-server,lmstudio,mock,ollama,openai,openai-compatible, andvllm.**kwargsAnyForwarded to the selected adapter constructor. Three factory-only keywords are also recognized:
metrics_profileconfigures model metrics after construction,metrics_enabledenables or disables their emission, andmax_parse_failuressets the validated consecutive action-parse failure limit without forwarding that ProtoLink-only option to the provider.
Returns
llmLLMAn initialized concrete adapter for the requested provider.
Raises
ValueErrorRaised when the provider string is unknown or
max_parse_failuresis outside the supported range.TypeErrorRaised when
max_parse_failuresis not an integer. Boolean values are not accepted as integers for this option.ImportErrorRaised when the selected adapter requires an optional dependency that is not installed.
provider or constructor errorCredential, model-path, URL, and client-construction errors from the selected adapter are not hidden.
Pass provider names as strings. Although the annotation includes LLMProvider, enum instances are not currently normalized correctly by the factory.
Examples
from protolink import LLMModelProfile, create_llm
llm = create_llm(
"openai-compatible",
base_url="http://localhost:1234/v1",
model="local-model",
metrics_profile=LLMModelProfile(context_window=32_768),
max_parse_failures=4,
)
max_parse_failures defaults to 3 and accepts integers from 1 through 10. Keep it as a top-level factory
argument. Do not place it in model_params, whose entries are provider generation options.
Base LLM contract
The LLM class defines the common interface that every implementation follows. Concrete adapters are deliberately thin at the runtime boundary: they translate ConversationHistory into a provider request, translate responses and streams back into ProtoLink values, and validate connectivity. The base class supplies the shared behavior around those calls.
Application code normally constructs a provider through create_llm() and calls chat() for direct text generation. Agent uses the more powerful infer() path, which adds typed actions, tool execution, delegation, policy checks, budgets, events, retries, and bounded iteration.
The core surface falls into four groups:
- Provider invocation:
call()andcall_stream()are implemented by concrete adapters. - Direct conversation:
chat()appends a user message and selects the blocking or streaming provider path. - Structured action acquisition:
call_action()andcall_action_stream()normalize native tool calls or JSON output intoLLMActionResult. - Runtime orchestration:
infer()repeatedly validates and dispatches those actions until the model returns a final answer.
ConversationHistory uses a collections.deque internally. Prepending or replacing the system prompt is an O(1) operation, and trimming older turns avoids repeated list reallocation on hot agent paths.
Do not instantiate LLM directly. Use a concrete implementation such as OpenAILLM, AnthropicLLM, OllamaLLM, or MockLLM, normally through create_llm().
LLM
class LLM(
model: str,
model_params: dict[str, Any],
*,
reasoning: Literal["none", "low", "medium", "high"] = "none",
)Abstract provider-neutral model contract. Subclasses implement text generation and connection validation; the base class owns history binding, typed action parsing, inference orchestration, compaction, and metrics.
Parameters
modelstrrequiredProvider-specific model identifier or local model path.
model_paramsdict[str, Any]requiredGeneration parameters passed to the adapter. The base class verifies only that later assignments remain dictionaries; the downstream SDK or server validates individual keys and values.
reasoning"none" | "low" | "medium" | "high"default: "none"Selects the reasoning instruction block included when the base system prompt is built. This argument is mainly for custom subclasses; current concrete provider constructors do not expose it.
Attributes
modelstrResolved model identifier.
model_paramsdict[str, Any]Mutable provider generation parameters.
historyConversationHistoryDefault history for direct calls, or the task-local history currently bound by
use_history().has_active_historyboolWhether the current execution context has a task-local history binding.
compactorHistoryCompactorComponent that mutates the live history using recent-message, token-budget, or summary compaction.
system_promptstrPrompt last built or assigned on the adapter.
metrics_profileLLMModelProfile | NoneOptional application-owned context-window and cost metadata.
metrics_enabledboolWhether inference may emit metrics when an observer is attached.
max_parse_failuresintMaximum consecutive JSON-action parsing or validation failures allowed during one inference run. Defaults to
3and accepts values from1through the ten-step inference ceiling. This runtime control is deliberately separate from provider generation parameters.syncSyncLLMBlocking wrapper for
infer(). Do not use it inside an active event loop.
The ten-step inference limit is protolink.llms.base.MAX_INFER_STEPS. It is a module constant, not an LLM class attribute.
LLM.chat
chat(
user_query: str,
*,
streaming: bool = False,
) -> str | AsyncIterator[str]Add one user message to the active history, then make a direct provider call.
Parameters
user_querystrrequiredUser text appended to the current
ConversationHistory.streamingbooldefault: FalseWhen
False, return the complete provider response. WhenTrue, return an asynchronous iterator of text chunks.
Returns
responsestrComplete model text when
streaming=False.chunksAsyncIterator[str]Asynchronous text stream when
streaming=True.
chat() appends the user message, but it does not append the returned assistant text. Use Agent state or manage ConversationHistory explicitly when you need durable multi-turn history.
Examples
answer = llm.chat("Give me a title for this report.")
async for chunk in llm.chat("Draft the introduction.", streaming=True):
print(chunk, end="")
LLM.call
call(
history: ConversationHistory,
) -> strGenerate one complete text response from a conversation history. Concrete adapters translate the history into their provider’s request format.
Parameters
historyConversationHistoryrequiredOrdered system, user, assistant, and tool messages sent to the model.
Returns
responsestrRaw text extracted from the provider response.
Raises
NotImplementedErrorRaised by the abstract base implementation.
provider errorConcrete adapters generally propagate authentication, HTTP, SDK, and model errors from the provider call.
LLM.call_stream
call_stream(
history: ConversationHistory,
) -> AsyncIterator[str]Start a streaming provider call and yield incremental text chunks.
Parameters
historyConversationHistoryrequiredConversation sent to the provider.
Yields
chunkstrThe next incremental piece of model output.
Consume the returned object with async for. Do not await call_stream() itself.
LLM.call_action
call_action(
history: ConversationHistory,
*,
tools: dict[str, BaseTool],
agent_callback_available: bool = False,
agent_cards: list[Any] | None = None,
) -> LLMActionResultAcquire one validated runtime action. The base implementation parses a JSON action from text; native-capable adapters override it and normalize provider tool calls into the same result type.
Parameters
historyConversationHistoryrequiredConversation for the current inference step.
toolsdict[str, BaseTool]requiredTool names mapped to executable tool objects. Native adapters turn this mapping into provider tool declarations.
agent_callback_availablebooldefault: FalseWhether the runtime can dispatch agent-delegation actions.
agent_cardslist[Any] | Nonedefault: NoneDiscovered agents available for native declarations and unambiguous fallback repair.
Returns
resultLLMActionResultA validated
FinalAction,ToolCallAction, orAgentCallAction, together with the raw response and provider metadata.
Raises
ValueErrorRaised when direct fallback parsing cannot produce a valid action.
LLM.call_action_stream
async call_action_stream(
history: ConversationHistory,
*,
tools: dict[str, BaseTool],
agent_callback_available: bool = False,
agent_cards: list[Any] | None = None,
chunk_callback: Callable[[str], Awaitable[None]] | None = None,
) -> LLMActionResultStreaming counterpart to call_action(). The fallback implementation forwards text chunks to the observer, buffers the complete response, and validates exactly one action after the stream ends.
Parameters
historyConversationHistoryrequiredConversation for the current inference step.
toolsdict[str, BaseTool]requiredTools available to the model.
agent_callback_availablebooldefault: FalseWhether agent delegation can be dispatched.
agent_cardslist[Any] | Nonedefault: NoneDiscovered agents available for delegation.
chunk_callbackCallable[[str], Awaitable[None]] | Nonedefault: NoneAsync observer invoked for each emitted text chunk.
Returns
resultLLMActionResultOne complete validated runtime action. Partial JSON fragments are never dispatched.
Raises
ValueErrorRaised when an empty or malformed fallback stream cannot be parsed as an action.
Controlled inference and tool use
What inference means in ProtoLink
infer() is the cornerstone of ProtoLink's agent runtime. A normal model call asks for text and returns text. Inference instead treats the model as a decision-maker inside a controlled loop: the model declares one typed action, ProtoLink validates and performs that action, the result is returned to the model as an observation, and the cycle continues until the model can produce a final answer.
This distinction is important. The LLM never executes Python code, invokes a remote agent, grants its own approval, or decides whether a budget may be exceeded. It can only request a final, tool_call, or agent_call action. ProtoLink remains the executor and policy boundary.
In normal applications, Agent prepares the prompt and calls infer() automatically. Direct calls remain available for advanced integrations that already have tools, callbacks, runtime policy, and history under their own control.
infer() enables a model to:
- Make tool calls - request external functions with structured arguments.
- Delegate to agents - pass work to a discovered specialized agent.
- Observe results - receive tool or agent output in conversation history.
- Self-correct - revise malformed actions, wrong arguments, or unavailable targets.
- Generate a final response - return user-facing content only after the required work is complete.
How action acquisition works
Every provider ultimately returns the same LLMActionResult, but it can acquire that action in one of two ways:
| Mode | Used by | Model instruction | Runtime behavior |
|---|---|---|---|
| JSON action mode | Default for local/small models and providers without reliable native tools | Return one JSON object such as {"type":"tool_call","tool":"search","args":{"q":"..."}} | call_action() or the fallback call_action_stream() parses text, validates it with Pydantic, and returns a typed action. |
| Native action mode | OpenAI, Anthropic, Gemini, and opted-in tool-capable servers | Use the provider's function or tool interface | The adapter sends real tool declarations, receives provider-native tool events, and normalizes them into the same typed action models. |
The rest of the loop is identical:
- Prompt selection: the Agent builds either the portable JSON prompt or the native-tool prompt. Streaming uses native instructions only when the adapter supports native streamed actions.
- Context preparation: the current query, history, tools, discovered agents, flow context, and runtime manifest are assembled.
- Action acquisition:
call_action()orcall_action_stream()obtains one model decision. - Validation: the decision becomes a
FinalAction,ToolCallAction, orAgentCallAction. Raw provider objects never reach the dispatcher. - Policy and budget checks: cancellation, authorization, capabilities, approvals, and remaining run budget are evaluated before side effects.
- Dispatch: ProtoLink executes a local tool, delegates to another agent, or accepts the final response.
- Observation injection: tool and agent results are appended through the provider-specific or provider-neutral history path.
- Iteration: the process repeats until
finalis produced or a guardrail stops the run.
Streaming JSON does not mean ProtoLink dispatches incomplete JSON fragments. The model streams ordinary text chunks that eventually form one complete action object. ProtoLink can forward those chunks to observers, but it buffers and validates the complete object before executing anything.
Ollama, llama.cpp, LM Studio, vLLM, and generic OpenAI-compatible servers default to JSON action mode. Enable supports_tool_calling=True only for a model and chat-template combination that reliably emits native tool calls.
Inference-loop safety guardrails
The inference loop includes multiple layers of protection against unreliable model output and runaway execution.
1. Deduplication detection
The runtime tracks a sliding window of successfully completed side-effect signatures. If the model requests an identical tool or agent action with identical arguments, ProtoLink:
- does not execute the duplicate;
- injects corrective guidance into history; and
- asks the model to use the existing observation, choose a different action, or finish.
Failed validation or execution does not poison the window, delegated infer signatures include the complete prompt, and final actions are returned immediately even when their text repeats. This prevents repeated side effects without suppressing a valid answer.
You have already performed this action. The result is in your context.
Proceed with the task: produce a final response or choose a different action.
2. Action parsing and the parse-failure circuit breaker
The parser is a security and interoperability boundary between untrusted model output and the runtime dispatcher. It
does not execute a tool or contact an Agent while trying to understand malformed text. It first produces one typed
FinalAction, ToolCallAction, or AgentCallAction; only that validated action can continue to policy, budget, and
dispatch checks.
Provider-native tool calls and portable JSON actions enter this boundary differently:
- A native-capable adapter translates the provider's structured tool event into the same Pydantic action models.
- A JSON-mode adapter passes the complete model text through the shared prompt-fallback parser.
The JSON path follows a deliberately conservative pipeline:
- Decode the untouched whole response. Valid JSON is never rewritten first. This preserves literal text such as
"<think>keep this</think>"inside a valid final response. - Recover syntax-only wrappers when whole-response decoding fails. ProtoLink can unwrap a complete JSON code
fence, ignore one complete leading
<think>or<thought>block, remove trailing commas outside JSON strings, or extract exactly one balanced object embedded in surrounding prose. The balanced scanner tracks nesting, quoted strings, and escapes in one pass. - Normalize only deterministic response shapes. A structured value placed in
FinalAction.contentcan be serialized losslessly as JSON text. A non-empty application object that omits the outer action envelope can become final content only when it contains no ProtoLink action-envelope fields. Existing legacy tool-call shorthands are repaired only when the tool and optional Agent target are unambiguous and their arguments are literal data. - Validate the complete action with Pydantic. Required fields, discriminated action types, forbidden extra fields, delegated-action combinations, and content types are checked before the infer loop sees the result.
Common small-model variations therefore have explicit outcomes:
| Model output | Parser behavior |
|---|---|
{"type":"final","content":{"answer":42}} | Serializes the object to JSON text and returns a FinalAction. |
{"answer":42,"sources":["doc-1"]} | Wraps the application object as final JSON content because no action-envelope field is present. |
```json {"type":"final","content":"done"} ``` | Unwraps the complete fence and validates the action. |
Result: {"type":"final","content":"done"} | Extracts and validates the single balanced object. |
{"type":"final","content":"done",} | Removes the unambiguous trailing comma outside strings, then validates. |
A complete leading <think>...</think> followed by one action | Ignores the reasoning wrapper only after untouched JSON decoding failed. |
The parser intentionally refuses cases where repair would require choosing or inventing meaning:
- two or more valid top-level JSON objects;
- an incomplete leading reasoning wrapper;
- an empty object or a non-object action payload;
- an explicit unknown action type;
- a tool- or Agent-shaped object with a missing action type;
- a missing tool, target, prompt, or argument that cannot be inferred uniquely; or
- an action found only inside a private reasoning wrapper with no public action after it.
Raw responses and parsed payloads use deterministic, 2,000-character head-and-tail previews in diagnostics. This keeps an oversized valid or invalid response from flooding logs and telemetry while retaining enough beginning/end context to diagnose truncation. For a decoded action that fails schema validation, correction history receives the concise field-level feedback and detected outer action type, not another copy of the parsed-payload preview.
Correction attempts
Parsing and action-schema failures are recoverable inside infer(). ProtoLink emits an llm_parse_error event,
retains bounded diagnostics for observability, and adds a correction message to conversation history. When JSON
decoding succeeded, the parser carries the decoded outer type separately from the rendered error. The retry can
therefore distinguish a malformed agent_call, malformed tool_call, invalid final, unknown action type, or object
with no recognizable outer type without scraping its own diagnostic text.
If the configured correction limit is exhausted, infer() raises InferParseError. It identifies the failed
inference step and number of consecutive attempts, gives a short explanation, keeps the final parser error, and
preserves the exact final model response. An empty Ollama response is reported as <empty> rather than appearing as a
blank diagnostic.
The correction also uses the current runtime capabilities. If the model selected agent_call but this inference has
no delegation callback, ProtoLink says that the action cannot be dispatched and does not show Agent-call examples. If
local tools or delegation are available, their canonical action shapes are included. ProtoLink does not silently
convert an explicit malformed side-effect request into final content.
Your previous response was not a dispatchable ProtoLink action.
Validation feedback:
Action validation failed. Field-level errors:
- Field 'agent_call -> action': Input should be 'tool_call' or 'infer'
The response selected `agent_call`, but no agent delegation route is available
for this inference, so that action cannot be dispatched.
Return exactly one JSON object using a currently dispatchable action:
- `final`: {"type":"final","content":"..."}
That is representative rather than a fixed template: field details and available alternatives depend on the rejected
payload, local tools, and delegation callback. A structurally valid tool_call or agent_call that reaches the
dispatcher but names an unavailable capability receives similarly explicit correction. It is not counted as a parse
failure because its action envelope was valid; the overall inference-step limit still prevents an endless correction
loop.
max_parse_failures controls the consecutive failure circuit breaker. It defaults to 3, accepts integers from 1
through 10, and resets after a successfully validated action:
from protolink import create_llm
llm = create_llm(
"ollama",
base_url="http://localhost:11434",
model="qwen3",
max_parse_failures=4,
)
Pass this option directly to create_llm(), or assign llm.max_parse_failures after construction. Do not put it in
model_params: those values are provider generation settings and are forwarded to the selected backend, while the
parse limit is ProtoLink runtime configuration.
The parse limit counts consecutive model action proposals that fail JSON decoding or action validation. It does not configure transient HTTP/provider retries, and it does not validate an application-specific schema inside final text. Those are separate boundaries whose retry budgets can multiply. Raising the limit can help a smaller model self-correct, but repeating a deterministic shape error is better handled by a lossless normalization or a clearer prompt than by unlimited retries.
from protolink import InferParseError
try:
result = await llm.infer(query="Review the evidence.", tools={})
except InferParseError as exc:
print(f"Failed after {exc.attempts} invalid proposals at step {exc.step}")
print(exc.explanation)
print(f"Last model response: {exc.raw_response!r}")
print(f"Parser detail: {exc.last_error}")
raw_response is the exact final provider text: "" means the provider returned no content, while None means an
adapter could not expose the text. The rendered exception keeps the existing bounded diagnostic preview, but the exact
attribute is not truncated. Treat it as potentially sensitive model output when logging or exporting it.
The action envelope is not the application schema
FinalAction.content is user-facing text. ProtoLink validates the outer action but cannot know whether that text must
also satisfy a domain-specific JSON schema. If an application expects a decision, invoice, ballot, or other structured
result, it should validate part.content separately and apply only domain-safe recovery:
from pydantic import BaseModel, Field
class Decision(BaseModel):
label: str
confidence: float = Field(ge=0, le=1)
part = await llm.infer(
query="Return the requested decision as JSON in final content.",
tools={},
)
decision = Decision.model_validate_json(part.content)
Do not infer missing high-impact fields from prose merely to make validation pass. A harmless public statement may permit a text fallback; a payment instruction, tool target, categorical decision, or authorization result normally should fail closed and request a new structured response.
A hallucinated tool call can be perfectly valid JSON with the correct field types. Parsing proves that the requested action is unambiguous and structurally valid; it does not prove that the action is factually correct, appropriate, or authorized. Keep tools narrowly allowlisted, use strict argument schemas and capability policies, require approval for important operations, enforce budgets and idempotency, and validate external identifiers at the execution boundary.
3. Self-correcting error recovery
Many model mistakes are observations, not immediate application failures:
| Error type | Runtime response |
|---|---|
| Unknown or unavailable tool | Lists tools that are actually available, or states that this inference has none. |
| Missing required fields | Returns field-level validation details. |
| Wrong tool arguments | Explains the callable mismatch and asks the model to check the input schema. |
| Agent not found | Reports the unavailable target through the delegation callback. |
| Agent delegation unavailable | States that no delegation route exists and suggests only dispatchable alternatives. |
| Invalid action type | Reports the detected type and lists only action kinds available in the current inference. |
Tool arguments are validated and conservatively coerced before authorization, deduplication, or tool-budget consumption. These model mistakes normally remain inside infer() so the model can correct them. An exception raised by the tool body, including TypeError, remains an execution failure and is not relabeled as an argument mismatch.
When an Agent-owned inference tool or delegation succeeds, the Agent also attaches and immediately snapshots an Artifact(kind="action_result") receipt keyed by the runtime action_id. The receipt records completion and correlation metadata but deliberately omits the internal result, which remains in private LLM history. If a later step fails, exceeds budget, or is canceled, the Task and optional RunStore snapshot still show what already happened without exposing model-only observations or inviting an unsafe blind retry. Non-JSON or circular results receive a bounded serialization fallback in that private history so the completion observation cannot be lost.
4. Physical retry safety
Transient provider failures use exponential backoff and jitter. Each physical attempt consumes the same LLM-call and input-token budget as an initial request and produces structured attempt metadata. Non-streaming calls may retry up to the configured retry limit; streaming calls may retry only before the first output chunk is exposed, preventing duplicated or discontinuous client output.
5. Bounded execution
MAX_INFER_STEPS limits a run to ten model decisions. If the model never produces a final action, the method raises RuntimeError with diagnostic context instead of looping indefinitely.
Inference event_callback failures are logged and the observer is disabled for the remainder of that infer call. A telemetry or UI exporter cannot turn an otherwise successful provider or tool operation into an application failure.
If “maximum inference steps exceeded” appears frequently:
- make the completion condition explicit in the system prompt;
- split a broad task into smaller agent or flow steps;
- improve tool names, descriptions, and input schemas;
- observe normalized inference events to inspect each decision; and
- verify that tool observations give the model enough information to finish.
Tool-call handling
Tool use has two separate phases:
- Action acquisition: the adapter obtains one model decision and returns
LLMActionResult. - Observation injection: after ProtoLink executes the tool, the adapter adds the result to history so the model can continue.
def call_action(...) -> LLMActionResult:
"""Return one validated action for the current inference step."""
async def call_action_stream(...) -> LLMActionResult:
"""Return one validated action from a streaming inference step."""
def _inject_tool_call(
self,
*,
tool_name: str,
tool_args: dict,
tool_result: Any,
) -> None:
"""Inject the runtime observation after a tool has executed."""
The base implementation asks for JSON, validates it, and injects a provider-neutral observation. Native adapters override action acquisition and observation injection where their API requires provider-specific tool-call IDs or message roles, but all paths converge before runtime dispatch.
Provider-specific action modes
| Provider | Non-streaming infer() | Streaming infer() | Notes |
|---|---|---|---|
| OpenAI | Native Responses function tools | Native streamed function-call events | Parallel tool calls are disabled so the runtime receives one action at a time. |
| Anthropic | Native tool_use blocks | Native streamed input_json_delta tool input | System instructions come from the task-local history; parallel tool calls are rejected. |
| Gemini | Native function declarations | Native streamed function-call parts | Function-call parts are normalized into ProtoLink actions. |
| DeepSeek | Native Chat Completions tools when supports_tool_calling=True | Native streamed tool deltas when enabled | Set the flag to False to force JSON action mode. |
| Grok | Native Chat Completions tools when supports_tool_calling=True | Native streamed tool deltas when enabled | Set the flag to False to force JSON action mode. |
| Ollama | JSON by default; native tools when explicitly enabled | JSON by default; native tool events when explicitly enabled | Keeps small/local model behavior simple unless tool support is known to work. |
| llama.cpp server/local | JSON by default; native tools when explicitly enabled | JSON by default; native tool events when explicitly enabled | Reliability depends on the selected model and chat template. |
| LM Studio / vLLM / OpenAI-compatible | JSON by default; native tools when explicitly enabled | JSON by default; native tool events when explicitly enabled | Dedicated subclasses provide conventional defaults and provider identity; the generic adapter covers LocalAI and compatible custom servers. |
| Hugging Face | JSON fallback for non-streaming inference | Not currently usable | call_stream() currently yields an empty chunk, so streaming inference is unsupported. |
Prompt selection
ProtoLink uses two prompt families:
- JSON prompt: describes
final,tool_call, andagent_callobjects. It is the compatibility path for small and local models. - Native prompt: tells the model to use the provider tool interface. It intentionally omits JSON action examples so native providers do not receive two conflicting tool protocols.
For streaming inference, prompt selection follows the adapter capability:
action_mode = "native" if llm.supports_native_action_stream else "json"
Native tool calls are not ordinary text. A provider may stream text deltas, function-argument fragments, or SDK objects. call_action_stream() hides those shapes and returns one complete typed action to the shared loop.
Agent delegation tools
Native providers receive synthetic delegation tools only when both conditions are true:
- the current runtime can dispatch agent calls; and
- discovered agent cards provide valid targets.
This avoids advertising a callable delegation interface that cannot succeed. JSON action mode can still produce an agent_call; if delegation is unavailable, the runtime injects corrective feedback instead of performing a side effect.
Design rationale
The layered design keeps the runtime strict without making every provider use the same wire protocol:
LLM.infer()dispatches one typed action at a time.- Provider adapters own provider-specific request and stream parsing.
- Small and local models keep a compact JSON protocol by default.
- On the Agent-prepared path, native providers use their real tool API without receiving JSON tool instructions.
- Every path converges on
FinalAction,ToolCallAction, andAgentCallAction. - Policy, approval, cancellation, budgets, retries, events, and execution remain outside the model.
Inference example
from protolink import Agent, AgentCard, create_llm
agent = Agent(
AgentCard(
name="weather-assistant",
description="Answers weather questions",
url="runtime://weather-assistant",
),
transport="runtime",
llm=create_llm("openai", model="gpt-4o-mini"),
)
@agent.tool(name="weather", description="Return the weather for a location")
async def weather(location: str) -> str:
return f"The weather in {location} is sunny."
answer = await agent.invoke("What's the weather in Tokyo?")
print(answer)
In JSON action mode, the intermediate model response must be exactly one supported object:
{
"type": "tool_call",
"tool": "weather",
"args": {"location": "Tokyo"}
}
After observing the tool result, the model completes with:
{
"type": "final",
"content": "The weather in Tokyo is sunny."
}
LLM.infer
async infer(
*,
query: str,
tools: dict[str, BaseTool],
agent_callback: Callable[[str, str, dict[str, Any]], Awaitable[Any]] | None = None,
agent_cards: list[Any] | None = None,
streaming: bool = False,
event_callback: Callable[[dict[str, Any]], Awaitable[None]] | None = None,
event_metrics: bool | None = None,
action_authorizer: Callable[[RunAction], Awaitable[ActionAuthorization]] | None = None,
cancellation_token: CancellationToken | None = None,
run_context: RunContext | dict[str, Any] | None = None,
budget_policy: BudgetPolicy | None = None,
budget_enforcer: BudgetEnforcer | None = None,
) -> PartRun the controlled multi-step inference loop used by Agent. The model declares typed intent; ProtoLink validates and executes tools or agent calls, feeds observations back to the model, and stops on a final action or safety limit.
Parameters
querystrrequiredUser task added to the active conversation history.
toolsdict[str, BaseTool]requiredExecutable tools available during this run.
agent_callbackCallable[[str, str, dict[str, Any]], Awaitable[Any]] | Nonedefault: NoneAsync dispatcher for delegated
tool_callandinferactions. Without it, attempted delegation is returned to the model as corrective feedback.agent_cardslist[Any] | Nonedefault: NoneDiscovered agents exposed to the model for delegation.
streamingbooldefault: FalseAcquire each model action from the streaming adapter path.
event_callbackCallable[[dict[str, Any]], Awaitable[None]] | Nonedefault: NoneAsync observer for normalized chunks, actions, tool events, delegation events, budget decisions, errors, and final output. The observer is non-authoritative; its first exception is logged and disables further callbacks for this infer call.
event_metricsbool | Nonedefault: NoneControls whether an attached event callback activates optional per-call metrics. Direct callers retain the existing default behavior; Agent sets this to
falsewhen its callback exists only to maintain internal completion receipts.action_authorizerCallable[[RunAction], Awaitable[ActionAuthorization]] | Nonedefault: NonePolicy boundary invoked after validation and before any tool or delegated-agent side effect.
cancellation_tokenCancellationToken | Nonedefault: NoneLive token checked before provider calls and action dispatch.
run_contextRunContext | dict[str, Any] | Nonedefault: NoneCorrelation, session, and budget context for the run.
budget_policyBudgetPolicy | Nonedefault: NoneAllow, warn, or deny policy used by the built-in budget enforcer.
budget_enforcerBudgetEnforcer | Nonedefault: NoneExisting stateful enforcer shared across several inference or tool operations. Omit it for an independent direct call; Agent supplies its task-local enforcer when the override accepts this keyword.
Returns
partPartA part with type
infer_outputand the final user-facing content.
Raises
InferParseErrorRaised when the configured number of consecutive model responses cannot be decoded or validated as ProtoLink actions. Exposes
attempts,step,explanation,raw_response, andlast_error.RuntimeErrorRaised after ten steps without a final action, or when an unrecoverable provider, tool, or delegated-agent failure is wrapped by the loop.
BudgetExceededErrorRaised when the configured run budget denies further work.
ApprovalRequiredError | ActionDeniedErrorRaised when runtime authorization requires approval or denies an action.
asyncio.CancelledErrorRaised when the run is cancelled.
Unknown tools, malformed actions, missing agents, and argument mismatches are normally reported back to the model for correction. They do not usually escape from infer() as ValueError.
Agent normally rebuilds the system prompt before calling this method. A custom runtime that invokes infer() directly must prepare matching tool and Agent descriptions with build_system_prompt(); supplying tools here only provides executables to the loop.
The new budget_enforcer parameter is optional. When an Agent uses a custom LLM.infer() override with the pre-0.6.7 signature, it supplies the shared enforcer only if the override declares that keyword or accepts arbitrary keyword arguments.
Examples
# Safe direct use when no tools or delegated Agents need to be advertised.
result = await llm.infer(
query="Summarize the supplied context.",
tools={},
)
print(result.content)
LLM.sync.infer
llm.sync.infer(
*,
query: str,
tools: dict[str, BaseTool],
agent_callback: Callable | None = None,
agent_cards: list[Any] | None = None,
streaming: bool = False,
event_callback: Callable | None = None,
) -> PartBlocking wrapper around LLM.infer() for scripts and synchronous command-line programs.
Parameters
querystrrequiredUser task passed to the asynchronous inference loop.
toolsdict[str, BaseTool]requiredTools available during inference.
agent_callbackCallable | Nonedefault: NoneOptional async agent dispatcher.
agent_cardslist[Any] | Nonedefault: NoneDiscovered agents available for delegation.
streamingbooldefault: FalseUse the underlying streaming action path.
event_callbackCallable | Nonedefault: NoneOptional async event observer.
Returns
partPartFinal
infer_outputpart.
This wrapper uses asyncio.run(). Do not call it from FastAPI handlers, asynchronous notebook cells, or any other active event loop; await llm.infer() there instead.
LLM.use_history
use_history(
history: ConversationHistory,
) -> ContextManager[ConversationHistory]Temporarily bind a conversation history to the current execution context. The binding uses contextvars, so concurrent asyncio tasks can share one LLM instance without sharing mutable messages.
Parameters
historyConversationHistoryrequiredHistory returned by
llm.historyinside the context.
Yields
historyConversationHistoryThe same object supplied by the caller.
Raises
TypeErrorRaised when
historyis not aConversationHistory.
Examples
from protolink.llms.history import ConversationHistory
customer_history = ConversationHistory("You support account customer-42.")
with llm.use_history(customer_history):
response = llm.chat("What do you remember about this account?")
LLM.compact_history
compact_history(
strategy: Literal["recent", "tokens", "summary"] = "recent",
*,
max_messages: int = 20,
max_tokens: int = 4000,
preserve_recent: int = 6,
summary_max_tokens: int = 512,
) -> HistoryCompactionResultCompact the active history in place. recent and tokens are local operations; summary makes one isolated synchronous model call before replacing older messages.
Parameters
strategy"recent" | "tokens" | "summary"default: "recent"recentretains a message window,tokensretains the newest suffix near a soft estimated-token ceiling, andsummaryreplaces older messages with model-generated durable context.max_messagesintdefault: 20Maximum retained messages for the
recentstrategy, including a leading system prompt.max_tokensintdefault: 4000Soft estimated-token ceiling for the
tokensstrategy. Protected messages may exceed it.preserve_recentintdefault: 6Newest non-system messages protected by
tokensandsummary.summary_max_tokensintdefault: 512Requested maximum length of a generated summary.
Returns
reportHistoryCompactionResultBefore/after message counts, estimated token counts, removed-message count, selected strategy, and whether a summary was created.
Raises
ValueErrorRaised for an unknown strategy, invalid limits, or an empty summary result.
provider errorErrors from the isolated summary call propagate when
strategy="summary".
llm.compactor.compact(...) has the same signature and behavior.
Examples
report = llm.compact_history("recent", max_messages=20)
report = llm.compact_history("tokens", max_tokens=8_000, preserve_recent=6)
report = llm.compact_history("summary", preserve_recent=8, summary_max_tokens=600)
print(report.to_dict())
LLM.configure_metrics
configure_metrics(
profile: LLMModelProfile | dict[str, Any] | None = None,
*,
context_window: int | None = None,
input_cost_per_million: float | None = None,
output_cost_per_million: float | None = None,
currency: str = "USD",
enabled: bool = True,
) -> LLMAttach application-owned model limits and prices used for observational inference metrics.
Parameters
profileLLMModelProfile | dict[str, Any] | Nonedefault: NoneComplete profile object or equivalent mapping. When supplied, it takes precedence and the individual context/cost arguments are ignored.
context_windowint | Nonedefault: NoneModel context window used to calculate pressure and remaining-token estimates.
input_cost_per_millionfloat | Nonedefault: NoneInput-token price per one million tokens.
output_cost_per_millionfloat | Nonedefault: NoneOutput-token price per one million tokens.
currencystrdefault: "USD"Currency label attached to calculated costs.
enabledbooldefault: TrueWhether inference may emit metrics events.
Returns
selfLLMThe same LLM instance for fluent configuration.
Metrics do not change provider payloads, retry behavior, or responses. They are emitted by infer() when an event observer or telemetry backend is attached.
Prompt architecture
ProtoLink keeps its prompt families in protolink/llms/prompts and chooses between them according to the action-acquisition mode. The runtime deliberately avoids giving a model two tool-calling contracts at once.
The system-prompt blueprint
LLM.build_system_prompt() assembles the prompt used by infer(). In JSON mode it describes the portable action objects. In native mode it tells the model to use the provider tool interface and leaves function-call syntax to the backend SDK or API.
By default, building a prompt resets history to the new system message. Passing persist=True updates the system message while retaining the rest of the current conversation, which is essential for long-lived sessions.
The final prompt is composed from:
-
Base instructions
- Define the model's role inside a deterministic runtime.
- Make it clear that the model requests actions rather than pretending to execute them.
- Select
BASE_SYSTEM_PROMPTfor JSON mode orNATIVE_SYSTEM_PROMPTfor native mode.
-
Tool instructions
- JSON mode injects the available tool descriptions and the
tool_callobject format. - Native mode adds a short instruction to use the provider tool interface; concrete schemas travel in the provider request.
- JSON mode injects the available tool descriptions and the
-
Agent capabilities
- JSON mode describes the
agent_callobject and discovered targets. - Native mode exposes synthetic delegation functions only when dispatch is available and agent cards exist.
- JSON mode describes the
-
Flow context
- Pipelines, routers, and graphs can inject topology-aware instructions for the current step.
- The model receives only the semantic context it needs for the active flow position.
-
Application instructions
- Your domain-specific prompt, such as “You are a coding assistant.”
- Appended to the shared runtime rules unless
override_system_prompt=True.
Tool and discovered-Agent metadata is serialized as deterministic valid JSON with stable ordering and explicit capabilities. The surrounding prompt labels those descriptions, schemas, and examples as untrusted data rather than executable instructions. This avoids Python-repr syntax and brace-escaping artifacts while keeping prompt caching and smaller-model parsing predictable.
Reasoning versus execution
When infer() runs, the prompt makes the LLM a reasoning and action-selection engine while ProtoLink remains the executor:
- Input: the model receives the task, history, tools, agents, and relevant flow context.
- Selection: it chooses
final,tool_call, oragent_call. - Structured output: it returns JSON or uses the provider-native tool channel.
- Validation: ProtoLink converts the result into a typed action.
- Execution: the runtime applies policy and performs the actual Python call or agent dispatch.
- Observation: the result is returned to the model for its next decision.
{
"type": "tool_call",
"tool": "get_weather",
"args": {"location": "Geneva"}
}
This separation of reasoning from execution is what allows one inference loop to support hosted APIs, local servers, and in-process models without handing runtime authority to untrusted model output.
LLM.build_system_prompt
build_system_prompt(
user_instructions: str | None = None,
agent_cards: str | None = None,
tools: str | None = None,
*,
action_mode: Literal["json", "native"] | None = None,
flow_instructions: str | None = None,
override_system_prompt: bool = False,
persist: bool = False,
agent_name: str | None = None,
) -> strBuild and store the complete runtime system prompt from base instructions, action mode, tools, discovered agents, flow context, and application instructions.
Parameters
user_instructionsstr | Nonedefault: NoneApplication instructions appended to the shared runtime prompt, or used as the complete prompt when
override_system_prompt=True.agent_cardsstr | Nonedefault: NoneSerialized discovered-agent descriptions used for delegation guidance.
toolsstr | Nonedefault: NoneSerialized tool descriptions used by the selected action prompt.
action_mode"json" | "native" | Nonedefault: NoneExplicit action acquisition mode. When omitted, the adapter’s
uses_native_action_promptproperty selects the mode.flow_instructionsstr | Nonedefault: NoneOptional pipeline, router, or graph context.
override_system_promptbooldefault: FalseReplace the shared runtime template with
user_instructions.persistbooldefault: FalsePreserve non-system history while updating the system message. The default clears history and resets it to the newly built system prompt.
agent_namestr | Nonedefault: NoneCurrent registered agent name used to prohibit self-delegation.
Returns
promptstrThe assembled prompt, also stored on
llm.system_prompt.
With persist=False, this method removes existing conversation turns. Use persist=True to retain them.
LLM.set_system_prompt
set_system_prompt(
system_prompt: str,
) -> NoneAssign a new value to llm.system_prompt.
Parameters
system_promptstrrequiredReplacement prompt string.
Returns
NoneNoneThis method mutates the adapter.
This setter does not rewrite the system message already stored in llm.history. Use build_system_prompt(..., override_system_prompt=True) when history and prompt state must be updated together.
LLM.validate_connection
validate_connection() -> boolCheck whether the provider client, server, or local model can respond.
Returns
connectedboolTruewhen adapter-specific validation succeeds; otherwise usuallyFalse.
Current concrete adapters already call validation during initialization. Most implementations catch validation failures, log them, and return False rather than failing construction.
API providers
Hosted-provider adapters read credentials from their conventional environment variable when api_key is omitted. They all implement direct call(), call_stream(), and validate_connection() methods and inherit chat(), history management, compaction, metrics, and the controlled infer() loop.
OpenAI, Anthropic, and Gemini always acquire actions through their provider-native function interface. DeepSeek and Grok use native Chat Completions tools by default but can be forced into portable JSON mode. Hugging Face currently supports non-streaming text generation only, so it is best suited to direct call() or chat(..., streaming=False) usage.
- OpenAI -
OpenAILLM, default modelgpt-4o-mini, credentialOPENAI_API_KEY. - Anthropic -
AnthropicLLM, default modelclaude-sonnet-4-20250514, credentialANTHROPIC_API_KEY. - Google Gemini -
GeminiLLM, default modelgemini-3-flash-preview, credentialGEMINI_API_KEY. - DeepSeek -
DeepSeekLLM, default modeldeepseek-chat, credentialDEEPSEEK_API_KEY. - Grok -
GrokLLM, default modelgrok-4-latest, credentialXAI_API_KEYorGROK_API_KEY. - Hugging Face -
HuggingFaceLLM, explicit model recommended, credentialHF_API_TOKEN.
OpenAILLM
class OpenAILLM(
*,
api_key: str | None = None,
model: str | None = None,
model_params: dict[str, Any] | None = None,
base_url: str | None = None,
)OpenAI Responses API adapter with native function tools and native streamed tool-call events. Use it for the official OpenAI service or for a custom base_url that implements the Responses API, not merely Chat Completions.
Direct calls translate ConversationHistory into Responses input and extract text from the returned response. In inference mode, real function declarations are sent to the provider, parallel tool calls are disabled, and returned function calls are normalized into ProtoLink actions before the runtime executes them. Streaming follows the same contract while forwarding text deltas and buffering function arguments until they form one complete action.
Parameters
api_keystr | Nonedefault: NoneOpenAI API key. Falls back to
OPENAI_API_KEY.modelstr | Nonedefault: NoneModel identifier.
Noneresolves togpt-4o-mini.model_paramsdict[str, Any] | Nonedefault: NoneValues merged over
temperature=1.0,top_p=1.0,top_logprobs=None, andtruncation="disabled".base_urlstr | Nonedefault: NoneOptional OpenAI client base URL. The endpoint must implement the Responses API; use
OpenAICompatibleLLMfor Chat Completions-only servers.
Raises
ImportErrorRaised when the OpenAI SDK is not installed.
OpenAI client errorMissing credentials and request errors originate from the SDK.
Examples
from protolink.llms.api import OpenAILLM
llm = OpenAILLM(
model="gpt-4o-mini",
model_params={"temperature": 0.3, "max_output_tokens": 800},
)
AnthropicLLM
class AnthropicLLM(
*,
api_key: str | None = None,
model: str | None = None,
model_params: dict[str, Any] | None = None,
base_url: str | None = None,
)Anthropic Messages API adapter with native tool_use actions and streamed tool-input deltas. The adapter derives both the separated system prompt and conversational messages from the supplied task-local ConversationHistory, converts ProtoLink tools to Anthropic tool schemas, and keeps the provider's tool-use identifier in action metadata.
After ProtoLink executes a requested tool, the adapter uses that identifier to inject the observation in the shape expected by the Messages API. Both streaming and non-streaming inference therefore share the same public LLMActionResult even though Anthropic's wire representation differs from OpenAI's. The runtime accepts one action per step and rejects multiple parallel tool_use blocks instead of silently selecting or merging them.
Parameters
api_keystr | Nonedefault: NoneAnthropic API key. Falls back to
ANTHROPIC_API_KEY.modelstr | Nonedefault: NoneClaude model identifier.
Noneresolves toclaude-sonnet-4-20250514.model_paramsdict[str, Any] | Nonedefault: NoneValues merged over
temperature=1.0,top_p=1.0, andmax_tokens=1024.base_urlstr | Nonedefault: NoneOptional Anthropic-compatible API base URL passed to the SDK.
Examples
from protolink.llms.api import AnthropicLLM
llm = AnthropicLLM(
model="claude-sonnet-4-20250514",
model_params={"max_tokens": 2048},
)
GeminiLLM
class GeminiLLM(
*,
api_key: str | None = None,
model: str | None = None,
model_params: dict[str, Any] | None = None,
base_url: str | None = None,
)Google GenAI adapter with native function declarations and native streamed actions. It converts conversation messages and tool schemas into Google GenAI content, generation configuration, and function declarations.
Function-call parts are normalized into ToolCallAction or AgentCallAction before dispatch. Text-only responses become final actions, so application and Agent code sees the same result types used by every other provider.
Parameters
api_keystr | Nonedefault: NoneGoogle API key. Falls back to
GEMINI_API_KEY.modelstr | Nonedefault: NoneGemini model identifier.
Noneresolves togemini-3-flash-preview.model_paramsdict[str, Any] | Nonedefault: NoneGeneration configuration merged over
temperature=1.0andtop_p=1.0.base_urlstr | Nonedefault: NoneStored by the shared API base class. The current Gemini client does not forward this value to
genai.Client.
Examples
from protolink.llms.api import GeminiLLM
llm = GeminiLLM(model="gemini-3-flash-preview")
DeepSeekLLM
class DeepSeekLLM(
*,
api_key: str | None = None,
model: str | None = None,
model_params: dict[str, Any] | None = None,
base_url: str | None = "https://api.deepseek.com",
supports_tool_calling: bool = True,
)DeepSeek Chat Completions adapter implemented through the OpenAI SDK with DeepSeek's API root. It supports ordinary text calls, incremental content streams, native Chat Completions tool calls, and streamed tool-argument deltas.
Native action acquisition is enabled by default. Set supports_tool_calling=False when the selected model behaves more reliably with ProtoLink's JSON action prompt; the surrounding inference loop, tool execution, and return types remain unchanged.
Parameters
api_keystr | Nonedefault: NoneDeepSeek API key. Falls back to
DEEPSEEK_API_KEY.modelstr | Nonedefault: NoneModel identifier.
Noneresolves todeepseek-chat.model_paramsdict[str, Any] | Nonedefault: NoneValues merged over
temperature=1.0andtop_p=1.0.base_urlstr | Nonedefault: "https://api.deepseek.com"DeepSeek-compatible API root.
supports_tool_callingbooldefault: TrueUse native Chat Completions tools. Set to
Falseto force the portable JSON action fallback.
Examples
from protolink.llms.api import DeepSeekLLM
llm = DeepSeekLLM(model="deepseek-chat")
GrokLLM
class GrokLLM(
*,
api_key: str | None = None,
model: str | None = None,
model_params: dict[str, Any] | None = None,
base_url: str | None = None,
supports_tool_calling: bool = True,
)xAI Chat Completions adapter using direct synchronous and asynchronous HTTP clients. It builds OpenAI-style message and tool payloads, parses content or tool calls, and normalizes usage metadata when the response includes it.
Native tools and streamed tool deltas are enabled by default. Disable supports_tool_calling to use the portable JSON action protocol with a model or endpoint that cannot reliably follow the native function format.
Parameters
api_keystr | Nonedefault: NonexAI key. Falls back to
XAI_API_KEY, thenGROK_API_KEY.modelstr | Nonedefault: NoneModel identifier.
Noneresolves togrok-4-latest.model_paramsdict[str, Any] | Nonedefault: NoneValues merged over
temperature=1.0.base_urlstr | Nonedefault: NoneAPI root.
Noneresolves tohttps://api.x.ai/v1.supports_tool_callingbooldefault: TrueUse native tool calls and streamed tool deltas. Set to
Falsefor JSON action mode.
Examples
from protolink.llms.api import GrokLLM
llm = GrokLLM(model="grok-4-latest")
HuggingFaceLLM
class HuggingFaceLLM(
*,
api_key: str | None = None,
model: str | None = None,
model_params: dict[str, Any] | None = None,
)Hugging Face Inference API adapter for non-streaming direct calls. It is useful when a hosted Hub model is available through the inference service and you want that model behind the same LLM interface.
Pass an explicit Hub model identifier. The current adapter does not implement a usable text stream or provider-native actions, and only part of model_params is forwarded by call(), so choose another adapter when streaming or agent tool loops are required.
Parameters
api_keystr | Nonedefault: NoneHugging Face token. Falls back to
HF_API_TOKEN.modelstr | Nonedefault: NoneHub model identifier. The effective built-in default is an empty string, so pass an explicit model for normal use.
model_paramsdict[str, Any] | Nonedefault: NoneValues merged over
max_new_tokens=512,temperature=1.0,top_p=1.0, andrepetition_penalty=1.0. The currentcall()path forwards onlytemperature.
HuggingFaceLLM.call_stream() is not implemented yet and currently yields one empty string. Do not use chat(..., streaming=True) or infer(streaming=True) with this adapter.
Examples
from protolink.llms.api import HuggingFaceLLM
llm = HuggingFaceLLM(model="your-org/your-chat-model")
Server providers
Server adapters connect to a model process over HTTP, whether that process runs on the same machine or on remote infrastructure. The server owns model loading and hardware resources; the ProtoLink adapter owns history serialization, request construction, streaming, action normalization, connection validation, and integration with the shared inference loop.
All server adapters inherit from ServerLLM. Their common configuration consists of a server URL, a model identifier understood by that server, optional generation parameters, and a supports_tool_calling capability flag. Native tool calling is opt-in because protocol compatibility alone does not guarantee that the selected model and chat template can use tools reliably.
The inherited model_params property can be replaced with a dictionary, set_system_prompt() updates the adapter's prompt value, and each concrete provider implements call(), call_stream(), and validate_connection() for its endpoint.
OllamaLLM
class OllamaLLM(
*,
base_url: str | None = None,
headers: dict[str, str] | None = None,
model: str | None = None,
model_params: dict[str, Any] | None = None,
supports_tool_calling: bool = False,
)Client for Ollama's /api/chat endpoint. It serializes ConversationHistory into Ollama messages and supports ordinary responses, streamed chunks, usage normalization, and optional native tool events.
JSON action mode is the default because local-model tool reliability depends on both the model and its template. Set supports_tool_calling=True only after verifying that the selected Ollama model produces correct native tool calls; direct chat() and call() usage does not require that flag.
Parameters
base_urlstr | Nonedefault: NoneOllama server root. Falls back to
OLLAMA_URL; if neither is supplied, construction raisesValueError.headersdict[str, str] | Nonedefault: NoneAccepted by the constructor, but the current request path does not forward custom headers.
modelstr | Nonedefault: NoneOllama model name.
Noneresolves togemma4:e4b.model_paramsdict[str, Any] | Nonedefault: NoneValues merged over
temperature=1.0,num_predict=8192, andnum_ctx=8192.supports_tool_callingbooldefault: FalseOpt into native Ollama tools. The default uses JSON action mode.
Raises
ValueErrorRaised for a missing base URL, unsupported URL scheme, missing hostname, or unavailable client during a call.
RuntimeErrorRaised for non-success Ollama responses or malformed response payloads.
Examples
from protolink.llms.server import OllamaLLM
llm = OllamaLLM(
base_url="http://localhost:11434",
model="qwen3",
)
LlamaCPPServerLLM
class LlamaCPPServerLLM(
*,
base_url: str | None = None,
headers: dict[str, str] | None = None,
model: str | None = None,
model_params: dict[str, Any] | None = None,
supports_tool_calling: bool = False,
)Direct client for a llama-server OpenAI-style Chat Completions endpoint. It talks to the server over HTTP without loading a model in the ProtoLink process, making it suitable when model lifecycle and hardware allocation belong to a separate service.
The adapter supports direct and streamed text calls. Native tool declarations are opt-in because correctness depends on the loaded model, chat template, and server build; JSON actions remain the compatibility default.
Parameters
base_urlstr | Nonedefault: NoneServer root. Resolution order is the argument,
LLAMACPP_SERVER_URL, thenhttp://localhost:8080.headersdict[str, str] | Nonedefault: NoneExtra HTTP headers.
Content-Type: application/jsonis added when absent.modelstr | Nonedefault: NoneServer model identifier.
Noneresolves togemma4:e4b.model_paramsdict[str, Any] | Nonedefault: NoneValues merged over
temperature=1.0.supports_tool_callingbooldefault: FalseOpt into native Chat Completions tools for a compatible model/template.
Examples
from protolink.llms.server import LlamaCPPServerLLM
llm = LlamaCPPServerLLM(base_url="http://localhost:8080")
OpenAICompatibleLLM
class OpenAICompatibleLLM(
*,
base_url: str | None = None,
api_key: str | None = None,
headers: dict[str, str] | None = None,
model: str | None = None,
model_params: dict[str, Any] | None = None,
supports_tool_calling: bool = False,
)Generic client for servers exposing /v1/chat/completions and /v1/models, including LocalAI and compatible custom services. Use it when the endpoint follows the Chat Completions protocol but is not the official OpenAI Responses API. Prefer VLLMLLM or LMStudioLLM when their conventional URL, credential environment variables, and provider identity are useful.
It supports custom headers, optional bearer authentication, direct and streamed content, and opt-in native tools. The default JSON response format makes the adapter well suited to ProtoLink's portable action protocol, while supports_tool_calling=True switches action acquisition to provider-style tool payloads.
Parameters
base_urlstr | Nonedefault: NoneServer root. Falls back to
OPENAI_COMPATIBLE_BASE_URL, thenhttp://localhost:1234/v1.api_keystr | Nonedefault: NoneOptional bearer token. Falls back to
OPENAI_COMPATIBLE_API_KEY.headersdict[str, str] | Nonedefault: NoneExtra request headers merged with the JSON defaults and authorization header.
modelstr | Nonedefault: NoneModel id passed to the server.
Noneresolves tolocal-model.model_paramsdict[str, Any] | Nonedefault: NoneValues merged over
temperature=1.0. Direct text calls also addresponse_format={"type": "json_object"}unless you override it.supports_tool_callingbooldefault: FalseEnable native Chat Completions tool payloads for a compatible server/model pair.
Examples
from protolink.llms.server import OpenAICompatibleLLM
llm = OpenAICompatibleLLM(
base_url="http://localhost:8000/v1",
model="Qwen/Qwen3-8B",
)
VLLMLLM
class VLLMLLM(
*,
model: str,
base_url: str | None = None,
api_key: str | None = None,
headers: dict[str, str] | None = None,
model_params: dict[str, Any] | None = None,
supports_tool_calling: bool = False,
)Convenience specialization of OpenAICompatibleLLM for a separately managed vLLM server. It inherits the compatible adapter's direct and streamed Chat Completions requests, portable JSON action fallback, usage normalization, connection validation, custom headers, optional bearer authentication, and opt-in native tools. The adapter itself does not install or import the vllm Python package.
The model is required because it must match the identifier accepted by the running server: normally the model passed to vllm serve, or a name configured with vLLM's --served-model-name option. The subclass supplies vLLM's conventional port and provider-specific environment variables and reports vllm in events and metrics.
Parameters
modelstrrequiredModel identifier exposed by the vLLM server. It must match the served model name.
base_urlstr | Nonedefault: NoneResolution order is the argument,
VLLM_URL, thenhttp://localhost:8000/v1.api_keystr | Nonedefault: NoneOptional bearer token. Falls back to
VLLM_API_KEY; no placeholder credential is added when both are absent.headersdict[str, str] | Nonedefault: NoneExtra request headers merged with the inherited JSON defaults and optional authorization header.
model_paramsdict[str, Any] | Nonedefault: NoneGeneration parameters forwarded to vLLM and merged over
temperature=1.0.supports_tool_callingbooldefault: FalseOpt into native Chat Completions tools only after the vLLM server and selected model are configured for automatic tool choice.
The default portable JSON action mode needs no vLLM tool parser. Before setting supports_tool_calling=True, start vLLM with --enable-auto-tool-choice and a model-appropriate --tool-call-parser, and ensure the selected model and chat template support tools.
Examples
from protolink.llms.server import VLLMLLM
# Start the model server separately: vllm serve Qwen/Qwen3-8B
llm = VLLMLLM(model="Qwen/Qwen3-8B")
LMStudioLLM
class LMStudioLLM(
*,
base_url: str | None = None,
api_key: str | None = None,
headers: dict[str, str] | None = None,
model: str | None = None,
model_params: dict[str, Any] | None = None,
supports_tool_calling: bool = False,
)Convenience specialization of OpenAICompatibleLLM for LM Studio. It keeps the complete compatible-server behavior while supplying LM Studio's conventional URL, credential fallback, and provider identity.
Use the generic parent class when you want environment variables and labels that are not tied to LM Studio. Use this subclass when local development should work with LM Studio's normal defaults and appear as lmstudio in events and metrics.
Parameters
base_urlstr | Nonedefault: NoneResolution order is the argument,
LMSTUDIO_URL, thenhttp://localhost:1234/v1.api_keystr | Nonedefault: NoneOptional token. Falls back to
LMSTUDIO_API_KEY, then the local placeholderlm-studio.headersdict[str, str] | Nonedefault: NoneExtra request headers.
modelstr | Nonedefault: NoneModel id selected in LM Studio. Inherited default:
local-model.model_paramsdict[str, Any] | Nonedefault: NoneGeneration parameters forwarded to the compatible server.
supports_tool_callingbooldefault: FalseOpt into native tools when the selected model supports them.
Examples
from protolink.llms.server import LMStudioLLM
llm = LMStudioLLM(model="local-model")
Local provider
Local adapters run inference inside the Python host rather than transmitting prompts to a server. This offers complete control over model files and data movement, but it also makes the application responsible for compatible native libraries, model loading, memory use, acceleration, and process stability.
LlamaCPPLocalLLM
class LlamaCPPLocalLLM(
*,
model: str,
model_params: dict[str, Any] | None = None,
supports_tool_calling: bool = False,
)In-process llama-cpp-python adapter for a local GGUF model file. Unlike LlamaCPPServerLLM, it loads the model in the current Python process, so model initialization time, native-library installation, memory use, and hardware configuration belong to the application.
The class implements complete and streamed chat-completion calls. Native tools are opt-in and depend on the loaded model and chat handler; JSON action mode is the safer default for portable inference.
Parameters
modelstrrequiredPath to the local model file.
model_paramsdict[str, Any] | Nonedefault: NoneValues merged over
temperature=0.8andmax_tokens=8192.supports_tool_callingbooldefault: FalseOpt into native llama.cpp tools when the loaded model and chat handler support them.
Raises
ImportErrorRaised when
llama-cpp-pythonis not installed.FileNotFoundErrorRaised when
modeldoes not exist.
The local package currently has no convenience init export. Prefer create_llm("llama.cpp-local", model="..."), or import the class from the full path shown above.
Examples
from protolink import create_llm
llm = create_llm(
"llama.cpp-local",
model="/models/qwen3-8b.gguf",
)
Testing provider
MockLLM
class MockLLM(
model: str = "mock-gpt",
model_params: dict[str, Any] | None = None,
*,
mock_responses: dict[str, Any] | None = None,
sequential_responses: list[Any] | None = None,
response_callback: Callable[[ConversationHistory, str], Any] | None = None,
default_response: str = "Unprocessed generic mock response",
)Dependency-free deterministic adapter for tests, examples, and offline runtime development. It implements the same LLM contract without network access, credentials, model files, or nondeterministic generation.
Responses are selected in a predictable priority order: a custom callback can inspect the full history, sequential responses can model multi-step action loops, keyword mappings can match prompts, and default_response handles everything else. This makes MockLLM suitable for testing Agent behavior, tool dispatch, parsing, history isolation, and failure paths rather than only simple chat.
Parameters
modelstrdefault: "mock-gpt"Identifier reported by the mock adapter.
model_paramsdict[str, Any] | Nonedefault: NoneOptional generation metadata retained on the instance.
mock_responsesdict[str, Any] | Nonedefault: NoneKeyword-based response mapping. Nested mappings can first match system-prompt text and then the latest user message.
sequential_responseslist[Any] | Nonedefault: NoneResponses consumed in order across calls.
response_callbackCallable[[ConversationHistory, str], Any] | Nonedefault: NoneCustom callback receiving the full history and system prompt.
default_responsestrdefault: "Unprocessed generic mock response"Fallback returned when no callback, sequential response, or mapping matches.
Examples
from protolink import create_llm
llm = create_llm(
"mock",
sequential_responses=[
{"type": "tool_call", "tool": "search", "args": {"query": "ProtoLink"}},
{"type": "final", "content": "Done"},
],
)
Related objects
LLMModelProfile
class LLMModelProfile(
context_window: int | None = None,
input_cost_per_million: float | None = None,
output_cost_per_million: float | None = None,
currency: str = "USD",
provider: str | None = None,
model: str | None = None,
supports_tools: bool | None = None,
supports_streaming: bool | None = None,
supports_json_schema: bool | None = None,
tokenizer: str | None = None,
metadata: dict[str, Any] = {},
)Immutable application-owned metadata used to calculate context pressure and estimated cost. It describes a deployment rather than configuring the provider request: changing this profile cannot enable tools, streaming, JSON schema support, or a larger model context window.
ProtoLink does not maintain a provider pricing catalog because limits and prices change independently of the library. Supply values from the provider contract used by your application and update them on your own release schedule. The class does not range-check those supplied values.
Parameters
context_windowint | Nonedefault: NoneTotal model context window in tokens.
input_cost_per_millionfloat | Nonedefault: NoneInput price per one million tokens.
output_cost_per_millionfloat | Nonedefault: NoneOutput price per one million tokens.
currencystrdefault: "USD"Currency label for estimates.
providerstr | Nonedefault: NoneOptional provider label.
modelstr | Nonedefault: NoneOptional model label.
supports_toolsbool | Nonedefault: NoneDescriptive tool support flag.
supports_streamingbool | Nonedefault: NoneDescriptive streaming support flag.
supports_json_schemabool | Nonedefault: NoneDescriptive JSON-schema support flag.
tokenizerstr | Nonedefault: NoneOptional tokenizer name used by application metadata.
metadatadict[str, Any]default: {}Additional application-defined metadata.
Usage examples
Basic chat
from protolink import create_llm
llm = create_llm("openai", model="gpt-4o-mini")
response = llm.chat("Hello, how are you?")
print(response)
Choose streaming as a separate interaction with a fresh or explicitly managed history:
streaming_llm = create_llm("openai", model="gpt-4o-mini")
async for chunk in streaming_llm.chat(
"Draft a short welcome message.",
streaming=True,
):
print(chunk, end="", flush=True)
chat() is deliberately small: it appends the user message and makes one provider call. It does not run tools, delegate to agents, apply the multi-step safety loop, or append the returned assistant text. Use infer() through an Agent when those runtime capabilities are needed.
Advanced inference with tools
import asyncio
from protolink import Agent, AgentCard, create_llm
async def main():
agent = Agent(
AgentCard(
name="calculator",
description="Performs checked calculations",
url="runtime://calculator",
),
transport="runtime",
llm=create_llm("openai", model="gpt-4o-mini"),
)
@agent.tool(name="multiply", description="Multiply two numbers")
async def multiply(left: float, right: float) -> float:
return left * right
answer = await agent.invoke("What is 15 multiplied by 8?")
print(f"Final answer: {answer}")
asyncio.run(main())
The Agent prepares the tool prompt and executable mapping, history binding, discovered-Agent context, policy boundary, cancellation token, run context, and budget configuration before invoking the LLM loop.
Updating parameters and prompts
llm.model_params = {
"temperature": 0.7,
"max_output_tokens": 500,
}
llm.set_system_prompt("You are a helpful coding assistant.")
Generation-parameter names are provider-specific. The base class requires a dictionary but does not translate keys such as max_tokens and max_output_tokens; the selected SDK or server decides which values are valid.
set_system_prompt() changes the adapter attribute only. To rebuild the runtime prompt and update the system message in history, use build_system_prompt(). Pass persist=True when existing turns must be retained.
Connection validation
if llm.validate_connection():
print("LLM is reachable.")
else:
print("LLM validation failed.")
Concrete constructors currently perform their own validation during initialization. Calling the method explicitly is still useful for health checks and diagnostics after a server, credential, network, or model state may have changed.
Error handling
LLM failures can originate at several different boundaries:
- Authentication errors: a provider rejects a missing, expired, or invalid credential.
- Connection errors: a server is unavailable, a URL is invalid, or the network call fails.
- Model errors: a model identifier is unknown, unavailable, or incompatible with the requested feature.
- Parameter errors: the downstream SDK or server rejects generation settings.
- Action errors: a model emits malformed or invalid structured intent.
- Execution errors: a selected tool, delegated agent, authorization policy, cancellation token, or run budget stops the loop.
- Parse guardrail errors: repeated parse failures produce
InferParseError. - Step guardrail errors: the ten-step safety limit produces
RuntimeError.
Recoverable action mistakes are normally injected back into history so the model can self-correct. Application-level exception handling should focus on provider failures and runtime boundaries that cannot be repaired inside the loop.
A direct call_action() or call_action_stream() invocation raises ValueError when its one response cannot become a
valid action. infer() catches that validation failure, emits llm_parse_error, requests a correction, and raises
InferParseError only when max_parse_failures consecutive proposals have failed. Inspect the final field-level
diagnostic before increasing the limit: repeated syntax drift may benefit from another attempt, while an unavailable
tool, ambiguous action, or application-schema mismatch needs a prompt, capability, or application-layer fix.
import asyncio
from protolink import (
ActionDeniedError,
ApprovalRequiredError,
BudgetExceededError,
InferParseError,
create_llm,
)
async def safe_inference():
llm = create_llm("openai", model="gpt-4o-mini")
try:
result = await llm.infer(
query="Summarize the available information.",
tools={},
)
print(f"Success: {result.content}")
except BudgetExceededError as exc:
print(f"Budget stopped the run: {exc}")
except (ApprovalRequiredError, ActionDeniedError) as exc:
print(f"Policy stopped the action: {exc}")
except InferParseError as exc:
print(f"Model action parsing failed: {exc.explanation}")
print(f"Last model response: {exc.raw_response!r}")
except asyncio.CancelledError:
print("The run was cancelled.")
raise
except RuntimeError as exc:
print(f"Inference failed: {exc}")
except Exception as exc:
print(f"Provider or configuration error: {exc}")
asyncio.run(safe_inference())
An in-process RuntimeTransport request still raises TransportRemoteError at the client boundary, while preserving
the typed infer error as its cause:
from protolink import InferParseError, TransportRemoteError
try:
task = await client.send_infer_task("Review the evidence.", agent_url)
except TransportRemoteError as exc:
if isinstance(exc.__cause__, InferParseError):
print(exc.__cause__.explanation)
print(f"Last model response: {exc.__cause__.raw_response!r}")
raise
Remote network transports serialize failures rather than sharing Python exception objects, so this direct cause access is specific to the in-process runtime transport.
Type aliases
The public type aliases describe provider-neutral model metadata:
from typing import Literal, TypeAlias
LLMType: TypeAlias = Literal["api", "local", "server"]
LLMProvider: TypeAlias = Literal[
"openai",
"anthropic",
"gemini",
"deepseek",
"grok",
"huggingface",
"llama.cpp-local",
"llama.cpp-server",
"lmstudio",
"mock",
"ollama",
"openai-compatible",
"vllm",
]
ReasoningLevel: TypeAlias = Literal["none", "low", "medium", "high"]
The factory also defines an LLMProvider enum with the same provider values. For current factory behavior, pass the string value such as "openai" or "ollama".
Migration guide
When migrating code written against earlier ProtoLink model wrappers:
- Replace
generate_response()withchat(). - Replace
generate_stream_response()withchat(..., streaming=True). - Use
Agent.invoke()for normal tool calling, delegation, policy, and multi-step execution. - Use
LLM.infer()directly only when implementing the surrounding prompt and runtime preparation yourself. - Await Agent or inference calls; direct non-streaming
chat()remains synchronous. - Expect a string from
chat()andAgent.invoke(), or aPartwith typeinfer_outputfrom directinfer().
# Earlier style
# response = llm.generate_response(messages)
# print(response.content)
# Direct text generation
response = llm.chat("Hello, how are you?")
print(response)
# Controlled Agent inference
answer = await agent.invoke("What's the weather?")
print(answer)
See also
- LLM examples for larger provider and tool-call examples.
- Agents for how
Agentbinds state and invokesLLM.infer(). - State for persistent conversation sessions.
- Runtime for cancellation, policy, approvals, budgets, and event recording.