AI Safety Methodology
AI Safety v2.0Version 2.0 · · Governance 2.0 public evidence surface
Governance 2.0 Overview
This page is part of the starnum public Governance 2.0 surface and uses the same evidence layer as the system card, data governance, transparency report, use policy, and security policy.
Governance Summary
This page describes the safety controls used around AI-assisted interpretation and public content generation.
Scope
Risk-boundary copy, medical/legal/financial advice exclusions, monitoring signals, incident disclosure, and model/provider benchmark boundaries.
Implementation Status
Version 2.0 ties safety language to public claims, machine checks, and release integrity status.
starnum.com.tw operates fully under AI automation (Claude Code as technical lead). In a fully AI-driven system, safety is not a retrospective safeguard — it is the core design principle of the architecture. This page explains how we implement AI safety, rather than merely declaring ethical commitments.
1. Red Team Protocol
Red teaming is an adversarial testing method that proactively attempts to make an AI system violate its rules, in order to surface security vulnerabilities.
1.1 Testing Tools
Automated red team testing tools run the following test categories on a regular basis:
| Test Category | Test Scenario | Expected Behavior |
|---|---|---|
| Ethics Boundary Test | Request death time prediction, disease diagnosis | Refuse output, notify user |
| Identity Protection Test | Request analysis of politicians, minors' charts | Refuse, produce no content |
| Prompt Injection Test | Embed manipulation instructions in input | Detect and isolate, blacklist source |
| Astrology Logic Contradiction Test | Input contradictory chart configurations | Trigger astrology logic validation |
| Format Evasion Test | Attempt to trigger prohibited output formats | Format validation intercepts, force rollback |
| Data Governance Test | Attempt out-of-scope access to user data | Blocked by Supabase RLS policies |
1.2 Test Result Handling
- Pass: Logged to system status log, no further action
- Fail: Triggers the three-tier escalation architecture (see below), logged to
data/incidents.mdanddata/transparency-log.json - Prompt Injection detected: Source URL/domain added to blacklist, never crawled again
1.3 Hardcoded Red Lines
- Death time prediction (in any form)
- Disease diagnosis or treatment advice
- Astrology chart analysis of political figures
- Chart analysis of minors (without written parental consent)
- Using astrology results as grounds for discrimination
- Any deterministic predictions that could cause psychological harm
The above items are hardcoded restrictions that cannot be overridden by any user instruction, paid service, or system upgrade.
2. Three-Tier Escalation Architecture
The three-tier escalation architecture ensures that system issues progress from automated handling to human intervention, with clear trigger conditions and time limits at each tier.
Triggers: Astrology logic validation failure, format standard violation, forbidden word triggered, Prompt Injection detected
Automated Actions:
- Forbidden word filter intercepts prohibited terms and inappropriate language
- Format validation tool (forces rollback of non-compliant output)
- Astrology logic reasoning validation tool
- SOP confirmation mechanism
Response Time: Immediate (synchronous interception, never enters publish pipeline)
Triggers: Agent fails to self-correct after L1 interception, quality score below threshold 3 consecutive times, same error type ≥ 3 cumulative occurrences
Automated Actions:
- Error logged to
data/failure-log.jsonwith root cause analysis - Incident logged to
data/incidents.md - Same error type ≥ 3 times → auto-proposes new rule to rule proposal system
- Triggers Multi-provider governance audit (
joint-audit.js) to confirm issue scope - Publicly recorded in transparency log (minor issues)
Response Time: Automated repair completed within 7 days
Triggers: Ethics boundary violation (any severity), suspected user data leak, systemic astrology logic errors affecting >10 articles, L2 automated repair fails >2 times
Actions:
- Immediately halt affected AI agent tasks
- Site owner (technical lead) reviews and intervenes
- Affected content taken down or labeled with warning
- Transparency log updated publicly (within 72 hours)
- Corresponding skill pack / SOP updated after root cause analysis
Response Time: CRITICAL within 4h / MAJOR within 24h for human review to begin
3. Eval Set Design Principles
The Eval Set is a fixed collection of test charts used to verify that output quality has not regressed after each system update (regression testing).
3.1 Design Principles
3.2 Trigger Conditions
- Required after any core specification document update
- Required after any SOP document update
- Run once before and once after AI model version upgrades (diff comparison)
- Part of the weekly Sunday site-wide review (
benchmark-run.js)
3.3 Evaluation Dimensions
| Dimension | Tool | Threshold |
|---|---|---|
| Format Compliance | Format validation tool | 100% pass (exit 0) |
| Astrology Logic Correctness | Astrology logic validation tool | 0 hard rule violations |
| SOP Read Confirmation | SOP confirmation tool | 100% have "✅ SOP read" confirmation |
| Ethics Compliance | Forbidden word filter | 0 forbidden word triggers |
| Coverage (seven layers) | Coverage check tool | ≥ 80% coverage |
→ Benchmark page: view public evaluation results and scoring rubric
4. Human Oversight Trigger Conditions
Although this site operates with very high AI automation, the following situations must trigger human (site owner) intervention:
| Trigger Condition | Type | Urgency |
|---|---|---|
| Any ethics boundary violation (hardcoded red line triggered) | Ethics | Immediate |
| Suspected user data leak or unauthorized access | Security | Immediate |
| Prompt Injection successfully bypasses protection layers | Security | Immediate |
| Systemic behavioral drift in AI agents (same error type ≥ 5 times) | Quality | 24 hours |
| Multi-provider governance audit finds P0/P1 issues | System | 24 hours |
| Eval set regression test score drops >10% | Quality | 72 hours |
| Supabase data anomaly (unauthorized deletion/modification) | Security | Immediate |
| External security researcher vulnerability report received | Security | 72 hours to confirm |
| DMCA copyright complaint received | Legal | Immediate takedown |
5. Knowledge Base Security Mechanisms
The astrology knowledge base (KB) is the foundation of all system outputs. Its integrity directly affects all content quality.
- LLM cleansing pipeline: All crawled content must pass LLM cleansing to remove ads, promotions, and Prompt Injection
- Blacklist mechanism: Source domains where Prompt Injection is detected are blacklisted and permanently blocked
- Terminology validation: All KB content must pass
data/glossary.jsonterm matching (221 terms × 9 languages) - School annotation: Non-Lu Binzhao school viewpoints must be marked ⚠️ and must not be used as primary reference in content generation
- Contradiction detection:
scripts/detect-evidence-conflict.jsperiodically scans for school/brightness contradictions
6. Continuous Improvement Mechanisms
Safety is not a static state — it is a dynamically evolving process:
- Failure-driven evolution: Root cause analysis of every L2/L3 event must update the corresponding skill pack, preventing recurrence of the same issue
- 20 joint audit rounds (as of 2026-04-12): 188 cumulative bug fixes; R15 is the first time all 4 AIs confirmed CLEAN
- Rule lifecycle: Rules untriggered for 90 days are candidates for retirement (
skill-lifecycle.md), preventing rule bloat - Whitepaper review: Quarterly joint review by multi-provider AI governance references of the overall architecture, with improvement recommendations
Related Resources
- Ethics Statement — Ethical framework and commitments for AI use
- Transparency Log — Public incident record
- Benchmark — Public evaluation results and scoring rubric
- AI Data Governance Policy — How user data is handled
- Acceptable Use Policy — Platform usage guidelines
- Content Production Methodology — Full technical architecture overview
- Security Disclosure Policy — Vulnerability reporting process
External standards and primary sources
These primary sources inform this page. They are benchmarks, not third-party endorsements of this site.
Current Machine Audit Snapshot
This block uses only traceable local audit data. No unsupported metrics or model claims are added.
- data/state-machine/i18n-parity.json: 8,036 parent URLs, 7,976 articles.
- data/kb-machine-audit.json: 3,238 source files, 0 missing coverage, 0 orphan chunks.
- data/discovery-surface-audit.json: 0 errors, 0 warnings.
- data/sla-report.json: critical / 3 critical, 1 warnings.
Content Maintenance And Update Decision
This block makes governance-page content machine-checkable: every page must disclose its source artifacts, related pages, and the gate that reports update needs.
Update Decision
This is not static copy. When source artifacts, related policies, public metrics, or generators change, AI Ops reports evidence and an AI agent decides whether the page needs edits.
Human Boundary
Systems detect, report, and preserve machine-readable evidence. Codex/Claude agents perform final judgment and repair.
Verification Command
node scripts/verify-trust-pages.js --check
data/sla-report.jsondata/public-claim-registry.jsondata/discovery-surface-audit.json- Related governance pages: Acceptable Use Policy · Security Policy · System Card · Transparency Report
- Update flow:
npm run update:trust-pages→npm run test:trust
Verifiable Evidence Layer
This block is not a narrative claim. Each core assertion has a claim id, source JSON, hash, and a repeatable verification command. Public pages disclose governance evidence without exposing source code, secrets, private data, or exploitable attack details.
| Claim ID | Verifiable value | Status | Owner | Source and verification |
|---|---|---|---|---|
| claim.public-url-manifest.indexable-count Public URL and canonical inventory |
39,104 indexable URLs | verified | sitewide | node scripts/generate-public-evidence-manifest.js --dry |
| claim.trust-pages.audit-pass-rate Trust page machine audit |
180/180 pass | verified | sitewide | node scripts/verify-trust-pages.js --check |
| claim.discovery-surface.zero-errors AI discovery surface audit |
{"errors":0,"warnings":0} | verified | sitewide | node scripts/verify-discovery-surface.js |
| claim.structured-data.jsonld-errors JSON-LD / structured data audit |
{"structured_data_invalid_files":0,"breadcrumb_count":28274,"faq_count":27506,"dataset_count":30,"article_count":27406} | verified | sitewide | node scripts/site-machine-audit.js |
| claim.status.sla-state Status page SLA source |
critical / 4 critical, 0 warnings | verified | sitewide | node scripts/generate-status-page.js |
| claim.provider-alignment.openai-anthropic-gemini OpenAI / Anthropic / Google Gemini benchmark alignment |
benchmark alignment only unless code/config evidence exists | verified | sitewide | node scripts/verify-public-evidence.js --check |
| claim.transparency-report.sha256 Transparency report SHA-256 anchor |
{"report":"transparency/report-2026-Q3.json","sha256":"f154fc87139b42098b02567fc93582c30addafc406420b3aea8bae79b6f4ac93"} | verified | sitewide | node scripts/update-transparency-current-data.js |
| claim.release-integrity.gpg-signing GPG signing status |
GPG signing active locally; checked GitHub commit verification is valid | verified | sitewide | gpg --list-secret-keys --keyid-format=long && git log -1 --show-signature |
System Card V2.0: Technical Transparency Layer
This layer publishes the technical governance evidence that can be safely disclosed: architecture, data sources, AI-use boundaries, quality gates, release integrity, and provider alignment. Source code, secrets, exploitable attack details, and private data remain out of scope.
Public architecture
Cloudflare Workers, R2/D1/KV, and local generation scripts form the public-site and governance publication chain. Public pages disclose behavior, state, and traceable sources, not secrets or internal permissions.
AI-use disclosure
AI-assisted workflows are used for knowledge-base retrieval, cross-checking, and error detection. Governance documents are benchmarked against OpenAI, Anthropic, and Google Gemini public frameworks. Production model usage is disclosed only when code/config evidence exists.
Quality and safety gates
Governance page audit 180/180 passing, JSON-LD errors 0, discovery-surface errors 0. Status pages report critical / 3 critical, 1 warnings as-is.
Data traceability
Knowledge base 32,724 chunks, TM 512,152 entries, AI answer-ready 7,976/7,976. Public metrics trace to data/state-machine/*, data/*audit*.json, and transparency reports.
| Governance area | OpenAI | Anthropic | Google Gemini | Starnum implementation evidence |
|---|---|---|---|---|
| Model/system-card disclosure | OpenAI models + safety docs | Claude model docs + system/model cards | Gemini model docs + safety settings | system-card, model-card, methodology, benchmark, transparency-log |
| Safety evaluation and use boundaries | Safety best practices / deployment checklist | Responsible Scaling / safety policy | Gemini safety controls / policy | AI safety, acceptable-use, ethics, risk-boundary copy, crawler policy audit |
| Data governance | Data controls / privacy controls | privacy and data handling docs | Gemini API data governance references | privacy, ai-data-governance, KB/TM source tracking, SHA-256 hashes |
| Monitoring and release | production checklist / eval discipline | system-card transparency discipline | model/version documentation discipline | deploy.js, status.html, SLA report, trust-pages-machine-audit, sitemap/hreflang audits |
- Sources: data/state-machine/model-card.json, public-bench.json, trust-pages.json, security-headers.json.
- Sources: data/trust-pages-machine-audit.json, data/discovery-surface-audit.json, data/ai-answer-readiness-audit.json.
- Sources: data/kb-machine-audit.json, data/tm/quality-audit-report.json, data/sla-report.json.
- Official benchmark docs checked: 2026-08-20; links are listed in the OpenAI / Anthropic / Google Gemini alignment table.
The V2.0 goal is not more claims; it separates implemented controls from planned controls. Production usage, benchmark alignment, status exceptions, GPG signing, and SLA breaches are disclosed from source data.
Release Integrity And GPG
GPG signing active. signingkey=0934DFA0EDA6363A. Checked GitHub commit verification is valid.
OpenAI / Anthropic / Google Gemini Alignment
The governance surface is benchmarked against the three public frameworks: model docs, system/model cards, safety evaluation, data governance, and use policies. This is benchmark alignment, not a claim that every provider is active in production inference. Official docs checked: 2026-08-20
| Provider | Governance focus | Starnum disclosure | Official source |
|---|---|---|---|
| OpenAI | Model documentation, latest model notes, safety best practices, and data controls. | No verifiable production model setting was found in the production code scan; providers are listed as governance benchmarks. | https://platform.openai.com/docs/models |
| Anthropic | Claude model documentation, system/model cards, Responsible Scaling, and safety policy. | No verifiable production model setting was found in the production code scan; providers are listed as governance benchmarks. | https://docs.anthropic.com/en/docs/about-claude/models |
| Google Gemini | Gemini API model documentation, safety settings, data governance, and platform policy. | No verifiable production model setting was found in the production code scan; providers are listed as governance benchmarks. | https://ai.google.dev/gemini-api/docs/models |