Weekly Update - June 15, 2026
Scoring bench bake-off, security audit, and three languages graduate from beta

This was a hardening-and-measurement week. Each of the three papers that launched under /papers the week before got its own dedicated blog post, marking a deliberate choice to give compliance officers, risk committees, and board-level readers a single-purpose entry point for each argument. Meanwhile, on a separate branch, the scoring benchmark was rebuilt from the ground up and a four-way model bake-off was run across Opus, DeepSeek, GPT-OSS, and Qwen. This is the Phase 2 instrument for evaluating whether the assessment's evaluator can migrate to open-weight models without sacrificing behavioral fidelity. Our quarterly full-repo security and documentation audit was conducted with the short-lived Fable 5 release, and produced findings with a sequenced mitigation plan already in motion. The multilingual site reached full translation coverage: every public page on PAICE.work is now available in Spanish, French, and Portuguese, with locale beta exit criteria defined and tooling built. And The Maturity Gap video named the organizational version of the measurement problem, that four teams managing four slices of AI risk, with none of them talking to each other.
Content Published Last Week
Monday (June 8): "Weekly Update - June 8, 2026"
Tuesday (June 9): "The Cost of Invisible AI Risk" Dedicated spotlight on the board-level business case paper. Your AI dashboards measure activity but none of them measure reliability. The paper frames AI risk exposure as three terms; the middle term (verification-failure rate) is the one no organization can currently compute. PAICE supplies it.
Wednesday (June 10): "The People-Vector Evidence Layer" Spotlight on the governance integration paper. Maps PAICE behavioral assessment output directly to NIST AI RMF, ISO/IEC 42001, and EU AI Act clauses. Designed so a compliance officer can lift the crosswalk table and place it directly into an audit file.
Thursday (June 11): "Governance Without Surveillance" Spotlight on the privacy architecture paper. Most behavioral measurement tools die in the approval room, not the evaluation room. The paper defines surveillance structurally — identification, retention, repurposability — and shows how PAICE fails all three tests by design.
Friday (June 12): Video - "The Maturity Gap" 3:22 video. Compliance works top-down. IT works bottom-up. Sales works outside-in. HR runs a survey. Four directions, all real, all disconnected. The maturity gap isn't the distance to the top of the curve — it's the distance between you and knowing where you are on it.
Papers Spotlight Series
Last week shipped the /papers library with three new papers alongside three legacy papers in a browsable, citable format. This week gave each new paper its own blog post — not a restatement of the paper, but a single-purpose entry point written for a different reader at a different stage of evaluation.
The Cost of Invisible AI Risk is for the CFO asking "what does this exposure actually cost?" The People-Vector Evidence Layer is for the compliance officer asking "where does this fit in my audit framework?" Governance Without Surveillance is for the works council representative or DPO asking "is this surveillance?"
Three papers, three objections, three dedicated posts. Each one designed to be the link someone forwards when the question comes up inside their organization.
Full Security Audit
A comprehensive security and documentation review landed on June 9, the same day as Anthropic's Mythos-class Fable 5 model release. We used this as an opportunity to test as well as conduct our scheduled quarterly audit a little early, which turned out to be a good thing as the model was only available for a few days before being revoked by federal order. Audit used automated scanning (Bandit, pip-audit, pnpm audit) plus LLM-driven OWASP Top 10 review via three parallel agents with independent verification of all findings.
The delta between this and our last security audit shows real progress on the remediation roadmap: the six SEC items fixed in May are confirmed resolved, and the new findings are either net-new or independent confirmation of items already tracked in the existing security remediation proposal.
Scoring Benchmark and Model Bake-off
Meanwhile, on the feature/scoring-bench branch (18 commits, not yet merged), the scoring benchmark was redesigned from the ground up and a four-way model bake-off was executed — a significant piece of work co-authored with Claude Fable 5 while Anthropic's research preview was available.
The scoring bench is the Phase 2 instrument for evaluating whether the PAICE assessment's evaluator — currently running on the Opus cascade — can migrate to open-weight models hosted on our own infrastructure. The design principle: behavioral ground truth is constructed, not labeled. Transcripts are generated or mutated with known catch/miss profiles, so Tier 1 truth is free. Invariants derived from the scoring philosophy (catch monotonicity, conservatism, false-alarm asymmetry) serve as an oracle-free comparison standard across candidate evaluators.
This branch was significantly advanced under Fable 5 availability, but was comporomized when it was revoked. Work continues to be delivered at a slower pace with Opus, and this branch is ready for L3 anchor labeling pass, after which the admission decision between Opus, DeepSeek, and GPT-OSS can be made on well-informed behavioral evidence.
Multilingual Site Completion
The i18n work that began with assessment flow translation in April reached full-site coverage this week. Every public-facing page — About, FAQ, Baseline, Accessibility, Homepage, Individual, Cohort, History, and system pages — is now translated in Latin American Spanish, French, and Brazilian Portuguese. That's 17 fully translated pages per language plus 809 UI strings covering every button, label, toast, error message, and navigation element. Full blog post announcement with details available this week.
Technical Improvements
Security Mitigations
Production database cleanup, private Github repo artifacts removed from the repository tree, and legacy archive files cleaned out. Gitignore patterns updated to prevent recurrence. The mitigation handoff documents the full finding-to-action map with effort estimates and decision gates.
Project Management Audit
Shipped tickets archived, new handoffs created. Packs pricing decisions recorded. Internal documentation index and URL inventory updated.
Platform Stability
All three environments (production, staging, pilot) operating normally. Zero unplanned incidents. Automated daily health checks and weekly backup restoration drills continue running without issue.
The Week in Numbers
- 5 posts published (3 paper spotlights + 1 video + 1 weekly update)
- 5 commits to main + 18 commits on
feature/scoring-bench(not yet merged) - 28 security findings documented (2 critical mitigated same-session)
- Scoring bench: 4 models evaluated, 52 L1 invariant pairs, 19 base transcripts, ~300 scoring runs
- Full-site i18n: 17 pages × 3 languages + 809 UI strings per locale
- Locale beta exit eval tooling shipped (7 automated gate checks)
- 1 production evaluator defect discovered via bench (narrative-numeric dissociation)
- 1 candidate model rejected on mechanical fitness (Qwen 3.5 397B)
- Security mitigation plan sequenced across 2 phases with decision gates
- 1 unplanned incident: manual PAICE Pro key support event
Why This Week Matters
The scoring benchmark is the most consequential piece of infrastructure that shipped this week, even though it hasn't merged yet. The PAICE assessment's behavioral fidelity depends entirely on the evaluator — the model that reads a transcript and produces dimensional scores. Right now that's Opus, which is excellent but expensive and not self-hosted. The new approach to the scoring bench gives us an objective, reproducible way to answer "can we replace it?" without guessing. Four models were swept against 52 invariant pairs and 19 base transcripts. The bench already caught a production evaluator issue that three months of live operation hadn't surfaced, and that alone justifies the construction cost. The bake-off ran during Anthropic's Fable 5 research preview window, and the collaboration with Fable 5 was instrumental in building the bench infrastructure at the speed required to complete the full sweep before the preview ended.
The papers spotlight series is a distribution investment. The papers already existed, they shipped the week before. This week's work was making each one findable on its own terms. A CFO Googling "AI collaboration risk cost" doesn't land on a papers index page. They land on a blog post that names their question in the title, makes the argument in five paragraphs, and links to the full paper for the team that needs the details. That's three entry points that didn't exist a week ago, each targeted at a different decision-maker in the same organization.
The security audit is worth naming publicly for the same reason we name it every quarter: organizations evaluating PAICE need to know that the platform they're trusting with behavioral assessment data is being audited regularly, that findings are documented with severity ratings and remediation plans, and that critical items don't sit in a backlog. Two critical findings were identified and resolved in the same session. The remaining items have a sequenced plan with clear ownership. That's what governance-grade operational hygiene looks like in practice.
The multilingual completion closes a gap that has been accumulating since April. A Spanish-speaking compliance director evaluating PAICE should be able to read the methodology, review the FAQ, and understand the Baseline engagement model without switching languages. That evaluation experience is now native-language end to end — from the first page through the assessment through the results.
Thank You
Thank you to everyone who read the papers spotlights this week and forwarded them to the person in their organization who needed each one. The Cost of Invisible AI Risk was written for board conversations. The People-Vector Evidence Layer was written for audit files. Governance Without Surveillance was written for works council meetings. If one of those links landed in the right inbox this week, the distribution model is working.
Get Involved:
- Take the assessment (free, always — ES, PT, FR available)
- Try Pulse (3-minute confidence check)
- Get PAICE Pro Packs (5 or 10 activations, team-ready)
- Browse the papers library (six papers, in-browser reader)
- Explore the AI Capability Baseline (cohort-level analytics for organizations)
- Contact us about your specific requirements
Related Reading
📖 This Week's Posts:
- The Cost of Invisible AI Risk — Board-level business case: the exposure your dashboards cannot see
- The People-Vector Evidence Layer — Governance crosswalk: PAICE → NIST AI RMF, ISO/IEC 42001, EU AI Act
- Governance Without Surveillance — Privacy architecture: why measurement is adoptable only when it isn't surveillance
- The Maturity Gap — Video: four teams, four slices of AI risk, zero shared picture
📖 Other Recent Updates:
- Weekly Update - June 8, 2026 — Papers library ships, dimensions complete, and PAICE enters a regulatory consultation
- Weekly Update - June 1, 2026 — Pro Packs launch, positioning refreshed, and the gap between confident and capable
Curious but short on time?
Take the 3-minute PAICE Pulse — a quick confidence check that maps how you see your own AI collaboration posture. No login required.