Methodology
Protocol v0.1.0 precommits categories, thresholds and metric definitions.

What security failures appear most often in AI-built applications? The protocol is public for inspection, but no result is shown until the dataset is large enough, confirmed, frozen, and independently reviewed.
The report is not ready to share as quantitative research. Its current evidence does not meet the precommitted sample, category-coverage, finding-confirmation, snapshot, or review requirements. Publishing charts now would turn repeated scans and unresolved candidates into unsupported prevalence claims.
Protocol v0.1.0 precommits categories, thresholds and metric definitions.
No frozen cohort exists. Exports must clear npm run research:dataset:audit before freeze.
Thresholds, Wilson intervals and gate evaluation are implemented and currently fail closed.
Chart components exist but render the gate-blocked state until validation passes.
An independent reviewer must reproduce every aggregate from the checked-in query.
The route stays noindex and out of the sitemap while any gate fails.
Distribution events cannot precede publication approval.
RLS, grants, policies, security-definer behavior, and object-level authorization.
Required evidence: rls_scan · supabase_advisors · repo_high_risk_read
Credential exposure in source, browser bundles, logs, or public environment variables.
Required evidence: code_static_scan · repo_high_risk_read · deployment_check
Missing identity verification, ownership checks, tenant boundaries, and privileged-route controls.
Required evidence: code_static_scan · repo_high_risk_read
Vulnerable, suspicious, or inadequately pinned packages in the resolved project dependency graph.
Required evidence: dependency_audit
Headers, storage exposure, CORS, debug surfaces, and production environment posture.
Required evidence: deployment_check
Prompt-injection boundaries, unsafe tool authority, SSRF paths, and missing execution stops or approvals.
Required evidence: llm_wiring_audit · deep_code_scan
Four approvals are required before any number is published. Each is recorded here with the requirement it certifies.
| Review | Requirement | State |
|---|---|---|
| Methodology review | Protocol, category definitions and thresholds precommitted before analysis. | Approved |
| Dataset review | Intake audit passes and exclusions reconcile against the frozen window. | Pending |
| Technical review | Independent reviewer reproduces every aggregate from the checked-in query. | Pending |
| Editorial review | Every claim traceable to a numerator, denominator and interval. | Pending |
These are chart contracts, not placeholder charts. No axes, percentages, or marks are rendered until a frozen dataset passes validation and every review is approved.
Render gate: At least 30 eligible projects in every displayed category
Render gate: A frozen dataset with documented exclusions
Render gate: Mutually exclusive project-level severity assignment and technical approval
Before publication, reviewers must decide whether platform-specific strata have enough coverage to report, whether repeat customers create correlated observations, and whether scanner-version changes require separate cohorts or a full backfill.