Evaluation harnesses · managed runtimes · retrieval quality · agentic security

Evaluation-first agentic AI production systems.

AIML Solutions builds production-grade evaluation harnesses and operating environments for teams moving from AI experiments to measurable, recoverable systems.

Agent regressionTrajectory diffingTool-call assertionsRAG evaluationMCPOpenClawOpenCodeDockerPythonFastAPICI gatesRuntime handoff

Hiring rather than buying? The operator's resume and public code are one click away.

Mission

Measure Before Shipping

Agent behavior, retrieval quality, tool use, cost, latency, and safety boundaries should be testable before a workflow is trusted.

Operate The Runtime

Useful agentic systems need scoped workspaces, recovery paths, approvals, and documentation, not just prompts.

Leave Evidence

Engagements produce harnesses, scorecards, runbooks, reports, or verification artifacts that can be reviewed after delivery.

MultiClaw OS is not a traditional operating system. It is an operating model and runtime pattern for coordinating agents, tools, data workflows, approvals, and evidence artifacts. The practical outcome is cleaner handoff, safer tool use, and more reproducible AI-assisted work.

Credibility And Current Proof

Terminal recording: triage-mesh assesses the seed repository, finds 46 live advisories, and stops at the human approval prompt
One command, real data. ./demo/run-demo.sh builds seven containers, pulls 46 live advisories from OSV.dev for the seed repo, and halts at the human gate. Zero token spend by default.
Jaeger UI showing one distributed trace spanning 7 services with 443 spans
One assessment, one trace. Jaeger showing a single distributed trace across all 7 services — 443 spans, depth 16 — from the orchestrator's request down to each policied tool call and OSV query.
triage-mesh architecture diagram: orchestrator, three A2A agents, three MCP servers, human gate, OSV egress, and the network-policy boundary
The boundary rule, drawn. A2A between orchestrator and agents; MCP from each agent to exactly one tool server; a single internet egress; the same policy file driving the harness, the compose networks, and the generated Kubernetes NetworkPolicies.

Live Runtime

IntelliClaw

Not a screenshot: this card is rendered from the latest cycle of IntelliClaw, an autonomous public-source signals pipeline that runs on a schedule with no API keys and no model spend. Harvest, normalize, cross-check, score, publish — the "operate the runtime" claim, observable.

signals processed in the latest cycle
scored high-risk by keyword and source-class confidence
public sources harvested
last cycle (UTC)
Top high-risk signals
  • Loading the latest cycle…
Flagship

Agent Regression Harnesses

Record/replay traces, tool-call assertions, trajectory diffs, cost/latency budgets, and CI gates for agent behavior drift.

Retrieval

RAG Quality Benchmarks

Domain-specific measurement rigs for citation grounding, context precision/recall, faithfulness, and abstention behavior.

Security

Agentic Safety Suites

MCP/tool-use authorization checks, prompt-injection fixtures, sandbox-boundary tests, and forbidden-action assertions. Reference implementation: triage-mesh.

Runtime

Managed Agent Workspaces

VPS, local, or hybrid OpenClaw/OpenCode-style environments with scoped workspaces, recovery paths, and handoff docs.

MultiClaw OS

Managed Runtime

Operating Model

MultiClaw OS is AIML Solutions' operating model for multi-agent work: scoped runtimes, orchestrated tool use, evidence artifacts, recovery paths, and human approval gates.

Managed Runtime

A productized setup for serious operators who want a working agentic AI command environment with documentation, boundaries, and support options.

Fleet Rollout

For agencies and teams that need multiple isolated workspaces for separate clients, projects, users, or research tracks.

Technical Credibility

Packt Technical Reviewer

Technical reviewer for OpenClaw AI in Production by Ken Huang, reviewing code, sequencing, reproducibility, dependencies, lifecycle, and security-boundary issues before publication.

Snorkel AI Evaluation Contributor

Contributes to frontier-model evaluation workflows involving reproducible environments, task setup, transcripts, tool-calling behavior, deterministic scoring, and failure recovery review.

Regulated Data Systems Background

15+ years of financial-risk, watchlist, entity-resolution, SQL, Python, and data-quality work inform the focus on provenance, evaluation, auditability, and operational discipline.

Public materials use sanitized examples. Private platform details, credentials, paid-task specifics, and client-sensitive information are excluded.

Open Proof Projects

Operator

Dennis Tien Donaghy

Agentic AI solutions engineer and technical reviewer based in Reno, NV. Operates hardened VPS-hosted OpenClaw/OpenCode-style multi-agent runtimes, contributes to Snorkel AI evaluation workflows, and brings 15+ years of regulated data engineering experience.

  • Packt technical reviewer: OpenClaw AI in Production
  • Snorkel AI evaluation contributor
  • Agentic AI Protocols: MCP, A2A, ACP
  • Open-source flagship: triage-mesh (MCP + A2A, red-team CI)
  • Kubernetes: KCNA/CKAD certification track (in progress)
  • UT Austin AI/ML for Business Applications

Engagement Process

Scope call and runtime/workflow inventory
Fixed deliverables and evidence targets
Implementation, audit, or evaluation review
Handoff docs and optional monthly support