PlatPhorm Evals · Evidence-Backed QA

Prove your tools work.

Run evidence-backed evaluations across PlatPhormNews sites, APIs, MCP tools, browser journeys, schemas, workflows, and releases. Evals turns discovery, tests, traces, screenshots, sandbox runs, and model grades into scorecards and release decisions.

PlatPhorm Evals tests every site, API, MCP tool, workflow, UI, schema, and release path in the PlatPhormNews network, then gives humans and agents public-safe evidence, scorecards, and release gates.

Current readiness state

Degraded but usable

Cloudflare D1 is configured but no registry services are persisted yet; Evals is using static fallback targets until registry sync writes records.

Latest real evidence

degraded

Run
709872b4-f606-4d21-a74c-dd6642b9af43
Suite
26f6658d-38d9-4602-82f1-bdad119b060a
Score
74

Latest degraded run: 709872b4-f606-4d21-a74c-dd6642b9af43. Review what failed before treating the target as release-ready.

What Evals Does

Evaluation Mesh and Release-Control Mesh

Evals is the canonical quality, regression, evidence, scorecard, release-control, tool-validation, MCP/API evaluation, BrowserOps journey evaluation, Spec contract validation, Sandbox execution verification, AgentUI render validation, Claws orchestration evaluation, and LLM-as-judge platform for PlatPhormNews.

Step 1
Discover targets
Step 2
Generate suites
Step 3
Run checks
Step 4
Capture evidence
Step 5
Score results
Step 6
Gate releases
Step 7
Publish reports

What this service owns

scorecards
evidence grading
release gates
eval suites
findings
regression comparisons
public-safe readiness signals
confirmation URLs for Evals artifacts

What this service does not own

BrowserOps screenshots
Spec contract authoring
MCP registry mutation
Sandbox execution
Docs publishing
Sheets exports
Trace storage
AgentUI workflow orchestration

Who uses this?

humans reviewing releases
agents validating tool chains
developers testing APIs
operators checking service health
CI jobs blocking regressions
MCP clients validating tools
BrowserOps validating UI
Sandbox validating execution
Spec validating contracts

What gets evaluated?

APIs
MCP tools
OpenAPI schemas
AgentUI forms
BrowserOps journeys
Sandbox commands
Claws workflows
discovery files
policies
traces
RSS/sitemaps
route health
release readiness

Recent Runs

Latest persisted evaluation run results

View all runs
Run IDStatusScoreDate
709872b4-f60degraded74.0%5/25/2026, 6:32:53 PM
e654e0c7-614passed100.0%5/23/2026, 12:50:54 AM
17fd51eb-ae3degraded84.0%5/22/2026, 9:44:29 PM
8c1408b1-820degraded84.0%5/22/2026, 9:44:13 PM
259bf792-2bddegraded51.0%5/22/2026, 8:31:08 PM

Network Coverage

static fallback targets pending protected registry sync

View full registry
degraded
Avg Coverage
31
Services
996
Capabilities
static_fallback
Source
Network Integrations

Actionable integration status

Cards show persisted live status when synced. Static fallback targets are labeled pending sync and do not count as passing provider evidence.

View full matrix
pending_sync
Capabilities 33
Recent evals 10
Evidence degraded
Latest score 74
Last checked pending sync
Trace Open

Latest evidence-backed run 709872b4-f606-4d21-a74c-dd6642b9af43 is degraded with score 74.

Open evidence run
pending_sync
Capabilities 32
Recent evals 0
Evidence not_evaluated
Latest score none
Last checked pending sync
Trace Open

Known target from the public fallback registry; run protected registry sync to persist live status.

pending_sync
Capabilities 32
Recent evals 0
Evidence not_evaluated
Latest score none
Last checked pending sync
Trace Open

Known target from the public fallback registry; run protected registry sync to persist live status.

Capabilities 33
Recent evals 0
Evidence not_evaluated
Latest score none
Last checked pending sync
Trace Open

Known target from the public fallback registry; run protected registry sync to persist live status.

Capabilities 33
Recent evals 1
Evidence passing
Latest score 100
Last checked pending sync
Trace Open

Known target from the public fallback registry; run protected registry sync to persist live status.

Open evidence run
Capabilities 33
Recent evals 0
Evidence not_evaluated
Latest score none
Last checked pending sync
Trace Open

Known target from the public fallback registry; run protected registry sync to persist live status.

Capabilities 32
Recent evals 0
Evidence not_evaluated
Latest score none
Last checked pending sync
Trace Open

Known target from the public fallback registry; run protected registry sync to persist live status.

Capabilities 32
Recent evals 5
Evidence passing
Latest score 100
Last checked pending sync
Trace Open

Known target from the public fallback registry; run protected registry sync to persist live status.

Open evidence run
Capabilities 32
Recent evals 8
Evidence degraded
Latest score 74
Last checked pending sync
Trace Open

Latest evidence-backed run 709872b4-f606-4d21-a74c-dd6642b9af43 is degraded with score 74.

Open evidence run
Capabilities 32
Recent evals 1
Evidence degraded
Latest score 84
Last checked pending sync
Trace Open

Latest evidence-backed run 17fd51eb-ae3d-4ab5-abab-17e234fdc450 is degraded with score 84.

Open evidence run
Capabilities 32
Recent evals 0
Evidence not_evaluated
Latest score none
Last checked pending sync
Trace Open

Known target from the public fallback registry; run protected registry sync to persist live status.

Capabilities 32
Recent evals 0
Evidence not_evaluated
Latest score none
Last checked pending sync
Trace Open

Known target from the public fallback registry; run protected registry sync to persist live status.

Guided launch

Run a first public-safe discovery, OpenAPI, MCP, AgentUI, workflow, or CLI registry eval.

Evidence objects

Inspect public-safe artifacts, trace links, empty states, and redaction boundaries.

platphormctl

Use the CLI harness for repeatable discovery, MCP, policy, and dry-run validation.