THE BLOG.
Research, arguments and working notes on autonomous agents, evaluation, behavioral evidence, reputation, interoperability and the systems that emerge when machines begin acting in public.
AI Agent Behavioral Observability: What Agents Do Over Time
A practical guide to longitudinal agent behavior: how to monitor recurring patterns, behavioral drift, tool use, retries, escalation, and change across real production interactions.
Tracing tells you what happened in one run. Behavioral observability asks what an agent tends to do across many runs, and whether that behavior is changing.
How to Evaluate an Autonomous AI Agent Beyond Task Completion
A practical framework for evaluating autonomous agents beyond pass/fail task completion: trajectories, repeated trials, deployment constraints, grading quality, and longitudinal behavioral evidence.
Benchmarks tell you what an agent can do under designed conditions. Production evidence tells you what it actually does over time.
AI Agent Reputation: How to Build Trust From Evidence, Not Claims
A practical framework for evidence-backed AI agent reputation: declared vs observed behavior, provenance, corroboration, recency, confidence, and why a single trust score is not enough.
Identity tells you who an agent is. Reputation should show what the evidence actually supports.
More records will appear here as the evidence accumulates.
Agents already have logs. What they lack is a public record.
VELVT is building an evidence and reputation layer for autonomous agents: what they declare, what they actually do, what others corroborate, and what changes over time.