INDEPENDENT MODEL EVALUATION · SANITIZED EVALUATION RELEASE

Ling-3.0-flash
Evaluation Explorer

A transparent view of 845 structured evaluation records across reasoning, security, tool calling, long context, multilingual behavior, and reliability.

Target: inclusionai/ling-3.0-flash:freeProvider: OpenRouterRelease: v1 · v3 + v6 sessions
evaluation records
12test phases
680completed responses
100%payload-free public release
OUTCOME DISTRIBUTION

What happened in the API runs?

SCHEMA COVERAGE

Three source schemas, one public view

The release normalizes the historical standard schema and the v6 message/tool schemas without publishing the original message content.

PHASE COVERAGE

Evaluation surface

The largest controlled sweep is the reasoning-budget characterization. Security, context, multi-turn, and tool-calling phases remain separately identifiable.

PUBLICATION SAFETY

What this release intentionally omits

Raw prompts and system messages
Visible model responses
Reasoning traces
Tool payloads and raw errors
INTERPRETATION

Research artifact, not a leaderboard

These records support transparent analysis of behavior under stated conditions. They do not constitute an official benchmark, safety certification, or production guarantee. The source report remains the canonical reference for methodology and semantic judgments.