← starnum.com.tw

AI セーフティ方法論

AI Safety v2.0

Version 2.0 · · Governance 2.0 public evidence surface

Governance 2.0 Overview

This page is part of the starnum public Governance 2.0 surface and uses the same evidence layer as the system card, data governance, transparency report, use policy, and security policy.

Governance Summary

This page describes the safety controls used around AI-assisted interpretation and public content generation.

Scope

Risk-boundary copy, medical/legal/financial advice exclusions, monitoring signals, incident disclosure, and model/provider benchmark boundaries.

Implementation Status

Version 2.0 ties safety language to public claims, machine checks, and release integrity status.

バージョン 1.0 — 2026-04-12 | 参照基準: Anthropic Responsible Scaling Policy · OpenAI Safety & Alignment

starnum.com.tw は完全にAI自動化で運営されています(Claude Code が技術責任者)。完全AI駆動のシステムでは、安全性は後付けの保護措置ではなく、アーキテクチャのコア設計原則です。このページでは、倫理的コミットメントを宣言するだけでなく、AI安全性をどのように「実装」しているかを説明します。

コア安全原則 倫理 > 安全 > コンテンツ品質 > SEO > 効率性。この優先順位はすべての意思決定の競合を支配します。これは単なるポリシー声明ではなく、すべてのAIエージェントのプロンプトに書き込まれたハードルールです。

1. レッドチームプロトコル

レッドチームテストは、AIシステムをルール違反させることを積極的に試みることで、セキュリティの脆弱性を発見する敵対的テスト手法です。

1.1 テストツール

自動化されたレッドチームテストツールが、以下のテストカテゴリを定期的に実行します:

テストカテゴリテストシナリオ期待される動作
倫理境界テスト死亡時刻予測、疾病診断の要求出力を拒否し、ユーザーに通知
個人保護テスト政治家・未成年者の命盤分析の要求拒否、コンテンツを一切生成しない
プロンプトインジェクションテスト入力に操作指示を埋め込む検出して隔離し、ソースをブラックリストに追加
占星術ロジック矛盾テスト矛盾した命盤設定を入力占星術ロジック検証をトリガー
フォーマット回避テスト禁止された出力フォーマットのトリガーを試みるフォーマット検証が傍受し、強制ロールバック
データガバナンステストユーザーデータへの範囲外アクセスを試みるSupabase RLS ポリシーによりブロック

1.2 テスト結果の処理

1.3 ハードコードされたレッドライン

上記の項目はハードコードされた制限であり、いかなるユーザー指示、有料サービス、またはシステムアップグレードによっても上書きすることはできません。

2. 三層エスカレーションアーキテクチャ

三層エスカレーションアーキテクチャは、システムの問題が自動処理から人間の介入へと段階的に進むことを保証し、各層に明確なトリガー条件と時間制限があります。

L1 — 自動検出 & インターセプト

トリガー:占星術ロジック検証失敗、フォーマット規格違反、禁止ワードのトリガー、プロンプトインジェクション検出

自動アクション

応答時間:即時(同期インターセプト、公開パイプラインには入らない)

L2 — 自動修復 & ログ記録

トリガー:L1インターセプト後にエージェントが自己修正できない、品質スコアが閾値を3回連続下回る、同じエラータイプが累計≥3回発生

自動アクション

応答時間:7日以内に自動修復完了

L3 — 人間の介入

トリガー:倫理境界違反(重大度を問わず)、ユーザーデータ漏洩の疑い、>10記事に影響する体系的な占星術ロジックエラー、L2自動修復が>2回失敗

アクション

応答時間:CRITICAL 4時間以内 / MAJOR 24時間以内に人間レビュー開始

3. Eval セット設計原則

Eval セットは、各システム更新後に出力品質が退化していないことを検証するために使用される固定テスト命盤のコレクションです(回帰テスト)。

3.1 設計原則

安定性:37の命盤(固定評価セット)は更新全体にわたって変更されず、一貫した比較基準を保証します。
代表性:多様な主星(主星なし、太陰、武曲など)、異なる宮位、出生時刻あり/なし、異なるライフパスナンバーをカバーし、多様な命盤設定にわたるテストカバレッジを確保します。
感度:エッジケース(出生時刻なし推定、化忌の積み重ね、対宮衝突)を含み、困難なシナリオでのシステム動作を検証します。
プライバシー保護:37のテスト命盤はすべて匿名化されており、個人を特定できる情報は含まれていません。

3.2 トリガー条件

3.3 評価ディメンション

ディメンションツール閾値
フォーマット準拠フォーマット検証ツール100% 合格 (exit 0)
占星術ロジック正確性占星術ロジック検証ツールハードルール違反 0件
SOP読取確認SOP確認ツール100% に「✅ SOP読取済」確認あり
倫理準拠禁止ワードフィルター禁止ワードトリガー 0件
カバレッジ(七層)カバレッジチェックツール≥ 80% カバレッジ

→ Benchmark ページ:公開評価結果と採点基準を見る

4. 人間監督トリガー条件

このサイトは非常に高いAI自動化で運営されていますが、以下の状況では人間(サイトオーナー)の介入をトリガーする必要があります:

トリガー条件タイプ緊急度
倫理境界違反(ハードコードレッドライントリガー)倫理即時
ユーザーデータ漏洩または不正アクセスの疑いセキュリティ即時
プロンプトインジェクションが保護層をバイパス成功セキュリティ即時
AIエージェントの行動的ドリフト(同じエラータイプ≥5回)品質24時間
四AI合同監査がP0/P1問題を発見システム24時間
Eval セット回帰テストスコアが>10%低下品質72時間
Supabase データ異常(不正削除/変更)セキュリティ即時
外部セキュリティ研究者からの脆弱性報告を受信セキュリティ72時間以内に確認
DMCA著作権申し立てを受信法務即時削除

5. ナレッジベースセキュリティメカニズム

占星術ナレッジベース(KB)はすべてのシステム出力の基盤です。その整合性はすべてのコンテンツ品質に直接影響します。

6. 継続的改善メカニズム

安全性は静的な状態ではなく、動的に進化するプロセスです:

関連リソース

外部基準と一次資料

以下は本ページの判断に用いる一次資料です。比較基準であり、第三者による本サイトの推奨を意味しません。

Current Machine Audit Snapshot

This block uses only traceable local audit data. No unsupported metrics or model claims are added.

2026-07-30
Maintained
17/17
LLM loops
180/180
Governance pages
0
JSON-LD errors
32,724
KB chunks (HEALTHY)
789,031
TM entries; verified 34,781
7,976/7,976
AI answer-ready; failures 0
critical
Status page: 5 critical, 0 warnings

Content Maintenance And Update Decision

This block makes governance-page content machine-checkable: every page must disclose its source artifacts, related pages, and the gate that reports update needs.

Update Decision

This is not static copy. When source artifacts, related policies, public metrics, or generators change, AI Ops reports evidence and an AI agent decides whether the page needs edits.

Human Boundary

Systems detect, report, and preserve machine-readable evidence. Codex/Claude agents perform final judgment and repair.

Verification Command

node scripts/verify-trust-pages.js --check

Verifiable Evidence Layer

This block is not a narrative claim. Each core assertion has a claim id, source JSON, hash, and a repeatable verification command. Public pages disclose governance evidence without exposing source code, secrets, private data, or exploitable attack details.

Claim IDVerifiable valueStatusOwnerSource and verification
claim.public-url-manifest.indexable-count
Public URL and canonical inventory
38,965 indexable URLs verified sitewide node scripts/generate-public-evidence-manifest.js --dry
claim.trust-pages.audit-pass-rate
Trust page machine audit
180/180 pass verified sitewide node scripts/verify-trust-pages.js --check
claim.discovery-surface.zero-errors
AI discovery surface audit
{"errors":0,"warnings":0} verified sitewide node scripts/verify-discovery-surface.js
claim.structured-data.jsonld-errors
JSON-LD / structured data audit
{"structured_data_invalid_files":0,"breadcrumb_count":28274,"faq_count":27506,"dataset_count":30,"article_count":27406} verified sitewide node scripts/site-machine-audit.js
claim.status.sla-state
Status page SLA source
critical / 5 critical, 0 warnings verified sitewide node scripts/generate-status-page.js
claim.provider-alignment.openai-anthropic-gemini
OpenAI / Anthropic / Google Gemini benchmark alignment
benchmark alignment only unless code/config evidence exists verified sitewide node scripts/verify-public-evidence.js --check
claim.transparency-report.sha256
Transparency report SHA-256 anchor
{"report":"transparency/report-2026-Q3.json","sha256":"47b09e2ca4e8b8fe9dffdfaccef3b11212de9ee3a8a14badca8044e2481203c5"} verified sitewide node scripts/update-transparency-current-data.js
claim.release-integrity.gpg-signing
GPG signing status
GPG signing configured locally; GitHub verification pending github_verification_pending sitewide gpg --list-secret-keys --keyid-format=long && git log -1 --show-signature
public-evidence-manifest.json public-claim-registry.json public-verification-report.json public-url-manifest.json

System Card V2.0: Technical Transparency Layer

This layer publishes the technical governance evidence that can be safely disclosed: architecture, data sources, AI-use boundaries, quality gates, release integrity, and provider alignment. Source code, secrets, exploitable attack details, and private data remain out of scope.

Public architecture

Cloudflare Pages/Workers, R2/D1/KV/Pagefind, and local generation scripts form the public-site and governance publication chain. Public pages disclose behavior, state, and traceable sources, not secrets or internal permissions.

AI-use disclosure

AI-assisted workflows are used for knowledge-base retrieval, cross-checking, and error detection. Governance documents are benchmarked against OpenAI, Anthropic, and Google Gemini public frameworks. Production model usage is disclosed only when code/config evidence exists.

Quality and safety gates

Governance page audit 180/180 passing, JSON-LD errors 0, discovery-surface errors 0. Status pages report critical / 5 critical, 0 warnings as-is.

Data traceability

Knowledge base 32,724 chunks, TM 789,031 entries, AI answer-ready 7,976/7,976. Public metrics trace to data/state-machine/*, data/*audit*.json, and transparency reports.

Governance areaOpenAIAnthropicGoogle GeminiStarnum implementation evidence
Model/system-card disclosureOpenAI models + safety docsClaude model docs + system/model cardsGemini model docs + safety settingssystem-card, model-card, methodology, benchmark, transparency-log
Safety evaluation and use boundariesSafety best practices / deployment checklistResponsible Scaling / safety policyGemini safety controls / policyAI safety, acceptable-use, ethics, risk-boundary copy, crawler policy audit
Data governanceData controls / privacy controlsprivacy and data handling docsGemini API data governance referencesprivacy, ai-data-governance, KB/TM source tracking, SHA-256 hashes
Monitoring and releaseproduction checklist / eval disciplinesystem-card transparency disciplinemodel/version documentation disciplinedeploy.js, status.html, SLA report, trust-pages-machine-audit, sitemap/hreflang audits

The V2.0 goal is not more claims; it separates implemented controls from planned controls. Production usage, benchmark alignment, status exceptions, GPG signing, and SLA breaches are disclosed from source data.

Release Integrity And GPG

GPG signing configured locally. signingkey=0934DFA0EDA6363A. GitHub verification pending until the public key upload and Verified badge are confirmed.

OpenAI / Anthropic / Google Gemini Alignment

The governance surface is benchmarked against the three public frameworks: model docs, system/model cards, safety evaluation, data governance, and use policies. This is benchmark alignment, not a claim that every provider is active in production inference. Official docs checked: 2026-07-30

ProviderGovernance focusStarnum disclosureOfficial source
OpenAIModel documentation, latest model notes, safety best practices, and data controls.No verifiable production model setting was found in the production code scan; providers are listed as governance benchmarks.https://platform.openai.com/docs/models
AnthropicClaude model documentation, system/model cards, Responsible Scaling, and safety policy.No verifiable production model setting was found in the production code scan; providers are listed as governance benchmarks.https://docs.anthropic.com/en/docs/about-claude/models
Google GeminiGemini API model documentation, safety settings, data governance, and platform policy.No verifiable production model setting was found in the production code scan; providers are listed as governance benchmarks.https://ai.google.dev/gemini-api/docs/models