Sentrail
Pre-publication · noindex

State of Vibe Coding Security 2026

What security failures appear most often in AI-built applications? The protocol is public for inspection, but no result is shown until the dataset is large enough, confirmed, frozen, and independently reviewed.

Technical summary

The report is not ready to share as quantitative research. Its current evidence does not meet the precommitted sample, category-coverage, finding-confirmation, snapshot, or review requirements. Publishing charts now would turn repeated scans and unresolved candidates into unsupported prevalence claims.

Publication pipeline

Methodology

Protocol v0.1.0 precommits categories, thresholds and metric definitions.

Dataset

No frozen cohort exists. Exports must clear npm run research:dataset:audit before freeze.

Validation

Thresholds, Wilson intervals and gate evaluation are implemented and currently fail closed.

Charts

Chart components exist but render the gate-blocked state until validation passes.

Technical review

An independent reviewer must reproduce every aggregate from the checked-in query.

Publication

The route stays noindex and out of the sitemap while any gate fails.

Distribution

Distribution events cannot precede publication approval.

Scope, dataset, and metric definitions

Population
AI-built applications voluntarily connected to Sentrail and meeting the report's evidence-completeness rules.
Unit of analysis
One eligible project at a frozen source revision
Prevalence
Eligible projects with at least one confirmed category finding divided by projects with complete evidence for that category.
Evidence state
Only confirmed findings enter a numerator. Candidate, rejected, suppressed, duplicate, and unresolved records are excluded.

Database and access control

RLS, grants, policies, security-definer behavior, and object-level authorization.

Required evidence: rls_scan · supabase_advisors · repo_high_risk_read

Secrets and environment handling

Credential exposure in source, browser bundles, logs, or public environment variables.

Required evidence: code_static_scan · repo_high_risk_read · deployment_check

Authentication and authorization

Missing identity verification, ownership checks, tenant boundaries, and privileged-route controls.

Required evidence: code_static_scan · repo_high_risk_read

Dependencies and supply chain

Vulnerable, suspicious, or inadequately pinned packages in the resolved project dependency graph.

Required evidence: dependency_audit

Deployment and configuration

Headers, storage exposure, CORS, debug surfaces, and production environment posture.

Required evidence: deployment_check

AI and agent tooling

Prompt-injection boundaries, unsafe tool authority, SSRF paths, and missing execution stops or approvals.

Required evidence: llm_wiring_audit · deep_code_scan

Methodology

  1. 01Freeze the study window, methodology version, scanner versions, source revision, and aggregation-query checksum.
  2. 02Select only the latest completed project audit in the window and assess eligibility independently for every category.
  3. 03Deduplicate at project and category grain; include only confirmed findings in numerators.
  4. 04Calculate category-specific prevalence with visible numerator, denominator, and Wilson 95% confidence interval.
  5. 05Reconcile exclusions, inspect included and excluded samples, obtain four review approvals, and freeze the snapshot before publication.

Review ledger

Four approvals are required before any number is published. Each is recorded here with the requirement it certifies.

ReviewRequirementState
Methodology reviewProtocol, category definitions and thresholds precommitted before analysis.Approved
Dataset reviewIntake audit passes and exclusions reconcile against the frozen window.Pending
Technical reviewIndependent reviewer reproduces every aggregate from the checked-in query.Pending
Editorial reviewEvery claim traceable to a numerator, denominator and interval.Pending

Visual evidence

These are chart contracts, not placeholder charts. No axes, percentages, or marks are rendered until a frozen dataset passes validation and every review is approved.

Failure prevalence by security category

Gate not met
Form
Comparison & Ranking · Horizontal bar with Wilson 95% confidence intervals
Denominator
Eligible projects with complete evidence for that category

Render gate: At least 30 eligible projects in every displayed category

  • ·No frozen cohort exists: the dataset export has not passed intake and freeze.
  • ·Technical review is not approved.

Evidence coverage by security category

Gate not met
Form
Comparison & Ranking · Horizontal bar
Denominator
All projects considered for the frozen snapshot

Render gate: A frozen dataset with documented exclusions

  • ·No frozen cohort exists: the dataset export has not passed intake and freeze.
  • ·Technical review is not approved.

Highest confirmed severity per affected project

Gate not met
Form
Composition · 100% stacked horizontal bar
Denominator
Eligible projects with at least one confirmed finding

Render gate: Mutually exclusive project-level severity assignment and technical approval

  • ·No frozen cohort exists: the dataset export has not passed intake and freeze.
  • ·Technical review is not approved.

Limitations, uncertainty, and robustness checks

  • The voluntary connected-app sample may differ materially from all AI-built applications.
  • Category denominators differ when a stack lacks an integration or a required check is incomplete.
  • Scanner versions, confirmation decisions, and repository revisions can change classification.
  • The study is descriptive and cannot establish that AI generation caused a security failure.
  • Sensitivity checks will compare definitions, exclusions, and highest-severity assignment before sign-off.

Required next steps

  • Collect enough independently eligible projects to clear overall and category thresholds.
  • Resolve or exclude every candidate finding in the analytical cohort.
  • Freeze and checksum the dataset; reproduce every aggregate from the checked-in query.
  • Complete methodology, data, technical, and editorial review.
  • Only then render charts, remove noindex, add the route to the sitemap, publish, and distribute.

Further questions

Before publication, reviewers must decide whether platform-specific strata have enough coverage to report, whether repeat customers create correlated observations, and whether scanner-version changes require separate cohorts or a full backfill.