命理分析品質ベンチマーク
Starnum Logic Engine v5.0 公開評価結果 | 37 件の固定命盤セット | 最終実行:
テスト概要
ベンチマークは、命理分析システムの品質を測定するための客観的な基準です。37件の固定命盤テストセット(ゴールデンテストスイート)を維持管理し、システム更新のたびに自動実行して分析品質が低下していないことを確認します。
テストセットは多様な命盤の組み合わせを網羅しています:14主星の異なる配置、12宮支、出生時刻あり・なし、異なる年干など — 実際のユーザーが遭遇しうるあらゆる状況を代表しています。
六次元評価基準(2026年4月)
多モデル交差検証加重合計スコア:85.4 / 100(シングルモデルベースライン79.2を上回り、+6.2向上)
| 次元 | 重み | 説明 | スコア |
|---|---|---|---|
| D1 | 30% | 正確性(四化 / 星位) | 87.3 |
| D2 | 20% | ルール完全性(格局 / 三方) | 82.1 |
| D3 | 20% | 解釈の深さ | 79.8 |
| D4 | 15% | 内部一貫性 | 91.2 |
| D5 | 10% | フック品質 | 84.6 |
| D6 | 5% | フォーマット準拠 | 96.4 |
データセットファイル(公開ダウンロード)
- 📊 dataset/classical_cases.json — 37件の匿名化命盤配置、命理フィールドのみを含み個人情報なし
- 📋 evaluation/scoring_rubric.json — D1–D6 六次元評価基準、詳細なルーブリックと検証方法を含む
- 📈 results/baseline-2026-04.json — 2026年4月ベースライン結果、多モデル加重スコア
- ⚙️ evaluate.js — 再現可能な採点・比較スクリプト(Node.js)
- 📄 README.md — 完全なドキュメントと使用説明
テスト方法論
テストセット設計
ゴールデンテストスイートは、異なる難易度レベルと星の組み合わせを代表する37件の固定命盤で構成されています:
- 主星の多様性:14主星すべてを網羅 — 紫微、天機、太陽、武曲、天同、廉貞、天府、太陰、貪狼、巨門、天相、天梁、七殺、破軍
- 宮位の多様性:命宮が12地支に分布し、すべての宮位の組み合わせを網羅
- 出生時刻あり・なし:既知の出生時刻(精密排盤)と未知の出生時刻(概算排盤)の両方を含む
- 年干カバレッジ:10天干すべてに代表的な事例があり、四化計算の正確性をテスト
評価基準
- 正確性:宮位の判断、四化の飛入、格局識別が陸斌兆派の基準に適合しているか — 127条のハードルールと照合
- 分析深度:白話文解釈が命宮主星の特徴、命身宮の関係、重要格局、個性、職業傾向などの核心面を網羅しているか
- 宮位カバー率:22の標準分析ブロックのうち実際に完成した割合
検証プロセス
- 固定テストセットをSupabase命盤データベースから読み込み、毎回同じ命盤を使用
- 命理ロジック検証ツールが127条のハードルールを自動照合し、潜在的エラーにフラグを立てる
- 多モデル交差比較:複数の独立したツールが同時に分析し、一貫性を相互確認
- 専門編集者による採点の人的審査、客観的で信頼性の高い評価結果を確保
- いずれかのルール検証が失敗した場合、バッチ全体を一時停止し修正後に再テスト
一般命理ウェブサイトとの比較
| 評価項目 | starnum.com.tw | 一般命理サイト |
|---|---|---|
| 四化派系の一貫性 | ✓ 全サイトで陸斌兆派に統一、明示的に注記 | ✗ 多派混用・説明なしが多い |
| 分析ロジック検証 | ✓ 127条ハードルールを自動照合 | ✗ 体系的な検証機構なし |
| 公開知識ソース | ✓ 8つの主要ソース、個別注記 | ✗ ソース不明または完全非公開 |
| 品質リグレッションテスト | ✓ 37件の固定命盤、バージョン更新後に自動テスト | ✗ 品質テスト機構なし |
| コンテンツ整合性検証 | ✓ SHA256 content hash + JSON-LD | ✗ 検証なし |
| 標準化された分析ブロック | ✓ 22の標準ブロック、一貫した構造 | ✗ 記事ごとに構造が異なり、深さも不均一 |
| 多モデル交差検証 | ✓ 複数ツールで交差比較 | ✗ 人間による単独レビュー |
比較基準は台湾と東南アジアの主要中国語命理サイトの公開コンテンツ。評価時期:2026年4月。
テストセット統計
| 分類 | 数量 | 説明 |
|---|---|---|
| 命盤総数 | 37 | 固定、バージョン更新後に差分をベースラインと比較 |
| 出生時刻あり | 32 | 精密な時宮計算が可能 |
| 出生時刻なし | 5 | 概算排盤処理能力をテスト |
| 年干カバレッジ | 8種 | 甲・乙・丙・丁・戊・己・庚・壬・癸すべてに代表事例 |
| 最多主星 | 太陰 | 命宮太陰の命盤がセット内で最も高い割合 |
| 生命数カバレッジ | 1–9 | 主要ライフパスナンバーすべてにテスト事例 |
このレポートについて
このページはStarnum Logic Engine v5.0の評価システムによって自動生成され、システムバージョンの更新ごとに再実行・更新されます。評価方法論は業界のソフトウェア品質保証(QA)基準を参照し、命理分析の特殊要件と組み合わせて設計されています。
私たちがこのベンチマークを公開することにしたのは、透明性こそが信頼を構築する唯一の方法だと信じているからです。評価方法論についてご質問がある場合は、Instagram @mychenanまでお気軽にご連絡ください。
外部基準と一次資料
以下は本ページの判断に用いる一次資料です。比較基準であり、第三者による本サイトの推奨を意味しません。
Current Machine Audit Snapshot
This block uses only traceable local audit data. No unsupported metrics or model claims are added.
- data/state-machine/i18n-parity.json: 8,036 parent URLs, 7,976 articles.
- data/kb-machine-audit.json: 3,238 source files, 0 missing coverage, 0 orphan chunks.
- data/discovery-surface-audit.json: 0 errors, 0 warnings.
- data/sla-report.json: critical / 5 critical, 0 warnings.
Verifiable Evidence Layer
This block is not a narrative claim. Each core assertion has a claim id, source JSON, hash, and a repeatable verification command. Public pages disclose governance evidence without exposing source code, secrets, private data, or exploitable attack details.
| Claim ID | Verifiable value | Status | Owner | Source and verification |
|---|---|---|---|---|
| claim.public-url-manifest.indexable-count Public URL and canonical inventory |
38,965 indexable URLs | verified | sitewide | node scripts/generate-public-evidence-manifest.js --dry |
| claim.trust-pages.audit-pass-rate Trust page machine audit |
180/180 pass | verified | sitewide | node scripts/verify-trust-pages.js --check |
| claim.discovery-surface.zero-errors AI discovery surface audit |
{"errors":0,"warnings":0} | verified | sitewide | node scripts/verify-discovery-surface.js |
| claim.structured-data.jsonld-errors JSON-LD / structured data audit |
{"structured_data_invalid_files":0,"breadcrumb_count":28274,"faq_count":27506,"dataset_count":30,"article_count":27406} | verified | sitewide | node scripts/site-machine-audit.js |
| claim.status.sla-state Status page SLA source |
critical / 5 critical, 0 warnings | verified | sitewide | node scripts/generate-status-page.js |
| claim.provider-alignment.openai-anthropic-gemini OpenAI / Anthropic / Google Gemini benchmark alignment |
benchmark alignment only unless code/config evidence exists | verified | sitewide | node scripts/verify-public-evidence.js --check |
| claim.transparency-report.sha256 Transparency report SHA-256 anchor |
{"report":"transparency/report-2026-Q3.json","sha256":"47b09e2ca4e8b8fe9dffdfaccef3b11212de9ee3a8a14badca8044e2481203c5"} | verified | sitewide | node scripts/update-transparency-current-data.js |
| claim.release-integrity.gpg-signing GPG signing status |
GPG signing configured locally; GitHub verification pending | github_verification_pending | sitewide | gpg --list-secret-keys --keyid-format=long && git log -1 --show-signature |
System Card V2.0: Technical Transparency Layer
This layer publishes the technical governance evidence that can be safely disclosed: architecture, data sources, AI-use boundaries, quality gates, release integrity, and provider alignment. Source code, secrets, exploitable attack details, and private data remain out of scope.
Public architecture
Cloudflare Pages/Workers, R2/D1/KV/Pagefind, and local generation scripts form the public-site and governance publication chain. Public pages disclose behavior, state, and traceable sources, not secrets or internal permissions.
AI-use disclosure
AI-assisted workflows are used for knowledge-base retrieval, cross-checking, and error detection. Governance documents are benchmarked against OpenAI, Anthropic, and Google Gemini public frameworks. Production model usage is disclosed only when code/config evidence exists.
Quality and safety gates
Governance page audit 180/180 passing, JSON-LD errors 0, discovery-surface errors 0. Status pages report critical / 5 critical, 0 warnings as-is.
Data traceability
Knowledge base 32,724 chunks, TM 789,031 entries, AI answer-ready 7,976/7,976. Public metrics trace to data/state-machine/*, data/*audit*.json, and transparency reports.
| Governance area | OpenAI | Anthropic | Google Gemini | Starnum implementation evidence |
|---|---|---|---|---|
| Model/system-card disclosure | OpenAI models + safety docs | Claude model docs + system/model cards | Gemini model docs + safety settings | system-card, model-card, methodology, benchmark, transparency-log |
| Safety evaluation and use boundaries | Safety best practices / deployment checklist | Responsible Scaling / safety policy | Gemini safety controls / policy | AI safety, acceptable-use, ethics, risk-boundary copy, crawler policy audit |
| Data governance | Data controls / privacy controls | privacy and data handling docs | Gemini API data governance references | privacy, ai-data-governance, KB/TM source tracking, SHA-256 hashes |
| Monitoring and release | production checklist / eval discipline | system-card transparency discipline | model/version documentation discipline | deploy.js, status.html, SLA report, trust-pages-machine-audit, sitemap/hreflang audits |
- Sources: data/state-machine/model-card.json, public-bench.json, trust-pages.json, security-headers.json.
- Sources: data/trust-pages-machine-audit.json, data/discovery-surface-audit.json, data/ai-answer-readiness-audit.json.
- Sources: data/kb-machine-audit.json, data/tm/quality-audit-report.json, data/sla-report.json.
- Official benchmark docs checked: 2026-07-30; links are listed in the OpenAI / Anthropic / Google Gemini alignment table.
The V2.0 goal is not more claims; it separates implemented controls from planned controls. Production usage, benchmark alignment, status exceptions, GPG signing, and SLA breaches are disclosed from source data.
Release Integrity And GPG
GPG signing configured locally. signingkey=0934DFA0EDA6363A. GitHub verification pending until the public key upload and Verified badge are confirmed.
OpenAI / Anthropic / Google Gemini Alignment
The governance surface is benchmarked against the three public frameworks: model docs, system/model cards, safety evaluation, data governance, and use policies. This is benchmark alignment, not a claim that every provider is active in production inference. Official docs checked: 2026-07-30
| Provider | Governance focus | Starnum disclosure | Official source |
|---|---|---|---|
| OpenAI | Model documentation, latest model notes, safety best practices, and data controls. | No verifiable production model setting was found in the production code scan; providers are listed as governance benchmarks. | https://platform.openai.com/docs/models |
| Anthropic | Claude model documentation, system/model cards, Responsible Scaling, and safety policy. | No verifiable production model setting was found in the production code scan; providers are listed as governance benchmarks. | https://docs.anthropic.com/en/docs/about-claude/models |
| Google Gemini | Gemini API model documentation, safety settings, data governance, and platform policy. | No verifiable production model setting was found in the production code scan; providers are listed as governance benchmarks. | https://ai.google.dev/gemini-api/docs/models |