Session lifecycle
SessionStarted, TurnStarted, TurnCompleted (with AgentFrame switch id where applicable).
Structured runtime records emit through TraceSink. The bundled JsonlTraceSink writes one record per line; fan-out, OTel, and custom sinks are wrappers. Lashlang execution graphs are a separate opt-in sink for foreground blocks, durable process runs, and trace-derived graph snapshots.
Attach a sink at builder time. It propagates to every session from that core.
use std::sync::Arc;
use lash::{
LashCore,
tracing::{JsonlTraceSink, TraceLevel, TraceSink},
};
let trace_sink: Arc<dyn TraceSink> = Arc::new(JsonlTraceSink::new("./.lash-data/trace.jsonl"));
let core = lash::LashCore::standard_builder(lash::TurnBudget::Unbounded)
.provider(provider)
.model(
lash::ModelSpec::builder(model.clone())
.context_window_tokens(200_000)
.build()
.expect("valid model metadata"),
)
.effect_host(Arc::new(lash::durability::InlineEffectHost::default()))
.attachment_store(Arc::new(lash::persistence::InMemoryAttachmentStore::new()))
.process_env_store(Arc::new(
lash::persistence::InMemoryProcessExecutionEnvStore::new(),
))
// Start bounded; tune both limits for your backend's latency envelope.
.commit_budget(lash::CommitBudget::bounded(1024 * 1024, 512))
.queued_work_batching(lash::QueuedWorkBatchingConfig::new(1024))
.trace_sink(trace_sink)
.trace_level(TraceLevel::Extended)
.build(crate::example_process_owner())?;
TraceLevel::Standard (default): turns, tool calls, LLM start/complete, token usage. TraceLevel::Extended adds provider stream chunks, prompt-component hashes, and runtime stream deltas. Extended is verbose; use for provider debugging.
Host-level Lashlang execution graphs are a separate opt-in sink. They are observability only: they are not written to process_events, do not create process-registry storage, and the normal trace_sink does not receive them unless the host explicitly tees the sink itself. Commands remain on the process admin surfaces: the session-scoped SessionProcessAdmin and the runtime-level LashCore::processes() handle; observation is optional trace projection.
let factory = lash::rlm::RlmProtocolPluginFactory::new(
lash::rlm::RlmProtocolPluginConfig::builder()
.instruction_limit(lash::rlm::InstructionBound::instructions(1_000_000))
.wall_clock(lash::rlm::WallClockBound::secs(30))
.memory_limit(lash::rlm::MemoryBound::mebibytes(64))
.build(),
std::sync::Arc::new(lash::persistence::InMemoryLashlangArtifactStore::new()),
)
.with_lashlang_execution_jsonl_path("./.lash-data/lashlang-execution.jsonl");
let core = lash::LashCore::rlm_builder(lash::TurnBudget::Unbounded, factory)
.provider(provider)
.model(model)
.effect_host(std::sync::Arc::new(
lash::durability::InlineEffectHost::default(),
))
.attachment_store(std::sync::Arc::new(
lash::persistence::InMemoryAttachmentStore::new(),
))
.process_env_store(std::sync::Arc::new(
lash::persistence::InMemoryProcessExecutionEnvStore::new(),
))
// Start bounded; tune both limits for your backend's latency envelope.
.commit_budget(lash::CommitBudget::bounded(1024 * 1024, 512))
.queued_work_batching(lash::QueuedWorkBatchingConfig::new(1024))
.build(crate::example_process_owner())?;
use std::sync::Arc;
use lash::tracing::{JsonlTraceSink, TeeTraceSink, TraceLashlangGraphStore, TraceSink};
let lashlang_graphs = Arc::new(TraceLashlangGraphStore::default());
let lashlang_execution_sink = Arc::new(TeeTraceSink::new([
Arc::clone(&lashlang_graphs) as Arc<dyn TraceSink>,
Arc::new(JsonlTraceSink::new("./.lash-data/lashlang-execution.jsonl"))
as Arc<dyn TraceSink>,
]));
let factory = lash::rlm::RlmProtocolPluginFactory::new(
lash::rlm::RlmProtocolPluginConfig::builder()
.instruction_limit(lash::rlm::InstructionBound::instructions(1_000_000))
.wall_clock(lash::rlm::WallClockBound::secs(30))
.memory_limit(lash::rlm::MemoryBound::mebibytes(64))
.build(),
Arc::new(lash::persistence::InMemoryLashlangArtifactStore::new()),
)
.with_lashlang_execution_sink(lashlang_execution_sink);
let core = lash::LashCore::rlm_builder(lash::TurnBudget::Unbounded, factory)
.provider(provider)
.model(model)
.effect_host(Arc::new(lash::durability::InlineEffectHost::default()))
.attachment_store(Arc::new(lash::persistence::InMemoryAttachmentStore::new()))
.process_env_store(Arc::new(
lash::persistence::InMemoryProcessExecutionEnvStore::new(),
))
// Start bounded; tune both limits for your backend's latency envelope.
.commit_budget(lash::CommitBudget::bounded(1024 * 1024, 512))
.queued_work_batching(lash::QueuedWorkBatchingConfig::new(1024))
.build(crate::example_process_owner())?;
let graph = lashlang_graphs.graph("process:process-id");
Tracking records use TraceEvent::LanguageExecution with typed payloads for execution start/finish, node start/complete/fail, branch selection, and child execution links. Producers stamp the language field as "lashlang". Every payload includes a deterministic event_key so TraceLashlangGraphStore dedupes replayed records. Graph keys are derived from runtime identity, such as effect:session:turn:effect-id for foreground code and process:process-id for durable processes. Graph snapshots are derived from trace events and safe for host UIs, dashboards, tests, and debugging; they are not canonical process state.
The v3→v4 rename also changed the OpenTelemetry span and attribute keys from lash.lashlang_execution* to lash.language_execution* and added lash.language_execution.language; update dashboards and alerts that key off those names.
When the linked Lashlang artifact contains static @label annotations, ExecutionStarted.execution_map.nodes carries each node's optional label_metadata with title and optional description. Lifecycle events continue to reference node_id only, and TraceLashlangGraphStore reduces the metadata from the initial execution map into graph nodes for host renderers.
Process wakes carry typed provenance through MessageOrigin::Process.caused_by. Hosts can inspect that semantic metadata to relate a wake back to a trigger occurrence, session node, or other runtime cause; rendering remains the host's responsibility.
A Lashlang graph is split into a static execution map and dynamic observation events. The static map gives a renderer the shape before any node runs; the dynamic stream updates node status, branch selection, timing, failures, and child execution links as the VM crosses observed sites.
ExecutionStarted carries TraceLanguageExecutionMap: graph identity, nodes, edges, node kinds, generated labels, and optional @label metadata. TraceLashlangGraphStore seeds those nodes as unobserved and branch edges as unknown. This is the data a UI uses to draw the full skeleton up front.NodeStarted, NodeCompleted, and NodeFailed update one node by node_id, set timestamps, duration, occurrence, and latest error. BranchSelected marks the chosen then or else edge as selected, marks the sibling branch edge as rejected, and completes the selected branch-arm node. ChildStarted links a parent node to another graph key, such as a started process graph. ExecutionFinished sets the graph-level status.@label@label(title: "...", description: "...") is static authoring metadata only. It travels in the execution map as label_metadata, becomes part of the module artifact identity, and never appears as a runtime command or dynamic update. A renderer should display the title/description from the static node and overlay the live status from dynamic events.process review_one(path: str) {
@label(title: "Review file")
report = await agents.default.spawn({
task: format("Review {}", path),
capability: "explore"
})?
finish report
}
@label(title: "Find candidates")
paths = await workspace.default.glob({ pattern: "src/**/*.rs" })?
@label(title: "Choose path")
if empty(paths) {
@label(title: "Submit empty result")
finish { reviewed: 0 }
} else {
@label(title: "Start child review")
child = start review_one(path: paths[0])
@label(title: "Collect child result")
report = (await child)?
finish { reviewed: 1, report: report }
}
For the run shown above, the graph snapshot keeps every static node even when the branch is not taken. A host UI can render the rejected branch faintly, keep the child-process link attached to the Start child review node, and update Collect child result from running to completed or failed when the next node event arrives.
{
"graph_key": "effect:session-1:turn-7:exec-1",
"status": "running",
"nodes": [
{
"id": "n1",
"kind": "operation",
"label": "workspace.default.glob",
"label_metadata": { "title": "Find candidates" },
"status": "completed",
"duration_ms": 12
},
{
"id": "n2",
"kind": "branch",
"label": "if empty(paths)",
"label_metadata": { "title": "Choose path" },
"status": "completed"
},
{
"id": "n3",
"kind": "start",
"label": "start review_one",
"label_metadata": { "title": "Start child review" },
"status": "completed"
}
],
"edges": [
{ "id": "e-then", "from": "n2", "to": "then", "label": "then", "selection": "rejected" },
{ "id": "e-else", "from": "n2", "to": "else", "label": "else", "selection": "selected" }
],
"children": [
{
"parent_node_id": "n3",
"child_graph_key": "process:review-process-id",
"child_entry_name": "review_one"
}
]
}
One synchronous append. Called on the runtime thread; must not block. JsonlTraceSink serializes and appends under a short-held mutex. A defaulted flush lets a host force buffered records to durable storage before process exit.
pub trait TraceSink: Send + Sync {
fn append(&self, record: &TraceRecord) -> Result<(), TraceSinkError>;
// Force buffered trace data to durable storage before exit.
// Default: no-op. Override when the sink buffers or can fsync.
fn flush(&self) -> Result<(), TraceSinkError> { Ok(()) }
}
Call flush before the process exits so records a sink has not yet committed are not lost. JsonlTraceSink::flush is honest about what it owns: each append already writes its record through to the OS (open, append, close — no in-process buffer), so flush only issues an fsync to push the OS page cache to disk. TeeTraceSink::flush fans out to every wrapped sink. StderrTraceSink and the default keep the no-op. A host that handed lash the sink already holds its own Arc and can flush it directly; LashCore::flush_trace_sink is the equivalent lever for hosts that did not retain the handle.
Fan out to multiple destinations by wrapping sinks:
struct FanoutTraceSink {
sinks: Vec<Arc<dyn TraceSink>>,
}
impl TraceSink for FanoutTraceSink {
fn append(&self, record: &TraceRecord) -> Result<(), TraceSinkError> {
for sink in &self.sinks {
// Treat errors per-sink; one failing destination shouldn't take the others down.
let _ = sink.append(record);
}
Ok(())
}
}
Tagged enums on TraceEvent. TraceEvent::kind() is the single source of truth for each variant's type tag — consumers match on the enum and read the kind from there rather than re-deriving tag strings. A new variant is breaking for closed-enum readers and bumps the schema; an optional field is additive only for readers that ignore unknown fields. The full rule set lives in Reporting channels → Schema evolution.
SessionStarted, TurnStarted, TurnCompleted (with AgentFrame switch id where applicable).
PromptBuilt: combined prompt_hash, prompt_chars, per-component fingerprints. CompositionChanged: one complete snapshot of the rendered system prompt and ordered model-facing tool schemas when that request composition differs from the last snapshot emitted by the resident session. Its SHA-256 key is built from cached rendered-prompt bytes and ordered cached tool-contract fingerprints, so prompt, membership, or same-member schema changes emit once while routing, model-capacity, and provider noise do not. An identical request performs no composition schema serialization and allocates no snapshot schema vector. A cold reopen is a new resident session, so its first request emits a fresh snapshot even when it matches the prior process's final composition. This is trace-only operational evidence, not a session observation shown to the model or user.
RollingHistoryCompactionNeeded and RollingHistoryPromptPruned are turn-scoped decisions; RollingHistoryCompactionStarted and RollingHistoryCompactionCompleted are runtime-operation-scoped because an explicit /compact can run without a turn. dropped_prefix_messages and retained_messages describe only the ephemeral prompt view; durable session history remains unchanged. A needed decision with no valid cut point terminates with a pruned record carrying zero dropped messages. Old attachment payloads can also be elided before the compaction threshold; that attachment-only optimization is deliberately untraced because the placeholder remains visible in the emitted provider request and no durable content changes.
LlmCallStarted, LlmCallCompleted, LlmCallFailed. LLM Provider, model, variant, request shape. LlmCallCompleted also carries generation_disposition when the adapter reports one: per generation option, whether the caller requested it and whether the request that ran actually carried it. ProviderReplayDropped records response-text, reasoning, or tool replay state rejected before LLM Provider serialization, including whether it was unstamped or minted by a different route and both the available minting route and selected serving route. This evidence is emitted at standard trace level on completion, failure, protocol abort, and cancellation; it does not depend on extended provider tracing.
ProviderRequest (the complete serialized request body with its byte length and SHA-256), ProviderStreamEvent (raw provider chunks), and RuntimeStreamEvent (post-projection SDK deltas).
ToolCallStarted, ToolCallCompleted. Args, output outcome (success / failure / cancelled), duration. Emitted per tool from one shared tool-execution seam, so a standard native call and a tool run inside a code block both produce exactly one Started + Completed pair. Containment (code block, batch parent) rides the matching TurnActivity; see Reporting channels.
TokenUsage: per-turn uncached input, output, cache-read input, cache-write input, and reasoning-output deltas. Reasoning is an output subset, not an additive total bucket. Mirrors the session usage ledger.
ProtocolStep: opaque host and plugin agentic-loop iterations. Runtime code execution uses the typed ExecCodeStarted, ExecCodeCompleted, ExecCodeFailed, and ObservationProjection events instead. ExecCodeCompleted carries a typed per-tool tool_calls roll-up (call_id, name, duration_ms, status); derive the count from the list.
Custom { name, payload }: escape hatch for host-specific events.
One line of UTF-8 JSON per record. TRACE_SCHEMA_VERSION = 9. Version 5 adds CompositionChanged; version 6 adds ProviderReplayDropped with typed minting and serving LLM Provider routes, including normalized endpoints; version 7 renames the Lashlang execution records to language-tagged ones; version 8 types the remaining turn/effect/wait/timer outcomes; version 9 promotes runtime exec diagnostics to four typed events, types the completed event's tool-call list, and removes the redundant tool_call_count and terminal_finish_present fields. Endpoints are routing metadata, not redacted secrets: URL userinfo is rejected, and hosts must not place credentials in paths or query strings because those bytes remain identity-significant and operator-visible in traces and remote envelopes. Typed decoding requires that exact version and returns TraceSchemaVersionError for any other version before interpreting the event payload. The schema-evolution policy distinguishes closed-enum variants from optional fields and opaque payloads.
Restate-hosted turns can attach the same sink to RestateRuntimeEffectController::with_trace_sink. The resulting records include journaled-effect start/settle, durable-wait park/resolve, durable timers, and segment boundaries alongside the ordinary session-manager turn, LLM, and tool records. The controller retains only a weak sink reference: an absent or dropped sink is a no-op. These records are live observations, may be emitted again during handler redrive, and never issue a Restate command or replace the durable journal.
Completed or failed LLM and tool events may also carry an additive attempts ladder projected from the retry owner's existing records. The records include the attempt count plus each attempt's outcome, duration, retry reason, and scheduled delay, so a throttling or transient-failure storm remains visible instead of collapsing into the final clean outcome.
{
"schema_version": 9,
"id": "6621dfa7-2cd0-4296-ac20-15c0f2d3cec1",
"timestamp": "2026-05-11T11:42:01.234+00:00",
"context": {
"session_id": "chat-123",
"turn_id": "turn-7"
},
"type": "tool_call_completed",
"call_id": "call-9",
"name": "fetch_url",
"args": { "url": "https://example.com" },
"output": { "outcome": { "status": "success", "payload": { "...": "..." } } },
"duration_ms": 8
}
Record type: lash_trace::TraceRecord. Event variants: lash_trace::TraceEvent. Both Deserialize; parse directly with serde_json::from_str::<TraceRecord>(). That documented decode path enforces the exact schema version structurally.
Optional otel-trace cargo feature converts events to OTel spans for export to Jaeger, Honeycomb, Datadog, or any OTLP backend. Off by default. The sink lives in lash-trace and is re-exported by lash-core under the feature.
# In your downstream Cargo.toml, enable OpenTelemetry on the supported facade.
lash-runtime = { version = "=0.1.0-alpha.113", features = ["otel-trace"] }
lash_core::OtelTraceSink wraps an opentelemetry tracer and converts every event to a span with payload attributes. Use in place of (or alongside) JsonlTraceSink:
use std::sync::Arc;
use lash::{
LashCore,
tracing::{OtelTraceSink, TraceLevel, TraceSink},
};
// Exporter/provider setup stays with the host; this reads the
// process-global OpenTelemetry tracer provider.
let sink: Arc<dyn TraceSink> = Arc::new(OtelTraceSink::from_global_provider());
let core = lash::LashCore::standard_builder(lash::TurnBudget::Unbounded)
.provider(provider)
.model(model)
.effect_host(Arc::new(lash::durability::InlineEffectHost::default()))
.attachment_store(Arc::new(lash::persistence::InMemoryAttachmentStore::new()))
.process_env_store(Arc::new(
lash::persistence::InMemoryProcessExecutionEnvStore::new(),
))
// Start bounded; tune both limits for your backend's latency envelope.
.commit_budget(lash::CommitBudget::bounded(1024 * 1024, 512))
.queued_work_batching(lash::QueuedWorkBatchingConfig::new(1024))
.trace_sink(sink)
.trace_level(TraceLevel::Extended)
.build(crate::example_process_owner())?;
OtelTraceSink::flush is deliberately a no-op: the sink only starts and ends spans on your tracer, while the buffering that risks span loss on exit lives in your BatchSpanProcessor / exporter. Flushing it is the host's duty — call force_flush() (or shutdown()) on your TracerProvider before the process exits. Lash never owns the provider, so it cannot do this for you; LashCore::flush_trace_sink flushes the sink only.
Span names are derived from the typed event. Turn, LLM, and tool lifecycles become nested spans (lash.turn, lash.llm, lash.tool); a tool produces one lash.tool span whether it was a standard native call or ran inside a code block, because both flow through the same emission seam. Composition snapshots become lash.composition.changed events with the fingerprint, prompt-character count, and tool count. Set include_payload_json only when the exporter should also receive the complete rendered prompt and ordered tool schemas. The typed ExecCodeStarted, ExecCodeCompleted, and ExecCodeFailed events collapse into a lash.exec_code family, with the precise event kind on the lash.protocol.diagnostic_phase attribute; ObservationProjection gets lash.observation_projection, and opaque host/plugin protocol steps stay lash.protocol_step. The channel map that ties these span names back to the other reporting surfaces is Reporting channels.