명리 분석 품질 기준 테스트
Starnum Logic Engine v5.0 공개 평가 결과 | 37개 고정 명반 세트 | 마지막 실행:
테스트 개요
기준 테스트(Benchmark)는 명리 분석 시스템의 품질을 측정하는 객관적 기준입니다. 우리는 37개의 고정 명반 테스트 세트(Golden Test Suite)를 유지하며, 각 시스템 업데이트 후 자동으로 실행하여 분석 품질이 저하되지 않았는지 확인합니다.
테스트 세트는 다양한 명반 조합을 포함합니다: 14개 주성의 다양한 배치, 12궁 지지, 출생 시간 유무, 다양한 연간 — 실제 사용자가 만날 수 있는 모든 상황을 대표합니다.
6차원 평가 기준 (2026년 4월)
다중 모델 교차 검증 가중 합계 점수: 85.4 / 100 (단일 모델 기준 79.2 초과, +6.2 개선)
| 차원 | 가중치 | 설명 | 점수 |
|---|---|---|---|
| D1 | 30% | 정확성 (사화 / 성위) | 87.3 |
| D2 | 20% | 규칙 완전성 (격국 / 삼방) | 82.1 |
| D3 | 20% | 해석 깊이 | 79.8 |
| D4 | 15% | 내부 일관성 | 91.2 |
| D5 | 10% | 훅 품질 | 84.6 |
| D6 | 5% | 형식 준수 | 96.4 |
데이터셋 파일 (공개 다운로드)
- 📊 dataset/classical_cases.json — 37개 익명화된 명반 구성, 명리 필드만 포함, 개인정보 없음
- 📋 evaluation/scoring_rubric.json — D1–D6 6차원 평가 기준, 세부 채점 기준 및 검증 방법 포함
- 📈 results/baseline-2026-04.json — 2026년 4월 기준 결과, 다중 모델 가중 점수
- ⚙️ evaluate.js — 재현 가능한 채점 및 비교 스크립트 (Node.js)
- 📄 README.md — 전체 문서 및 사용 안내
테스트 방법론
테스트 세트 설계
Golden Test Suite(황금 테스트 세트)는 다양한 난이도 등급과 성요 조합을 대표하는 37개의 고정 명반으로 구성됩니다:
- 성요 다양성: 자미, 천기, 태양, 무곡, 천동, 염정, 천부, 태음, 탐랑, 거문, 천상, 천량, 칠살, 파군 14개 주성 전부 포함
- 궁위 다양성: 명궁이 12지지 전반에 분포하여 모든 궁위 조합이 포함됨
- 출생 시간 유무: 출생 시간 있음(정확한 명반 작성)과 없음(근사 명반 작성) 두 가지 경우 포함
- 연간 커버리지: 10개 천간 모두 대표 사례 보유, 사화 계산 정확성 테스트
채점 기준
- 정확성: 궁위 판단, 사화 비입, 격국 식별이 육빈조파 기준에 부합하는지 — 127개 하드 규칙 전체 대조
- 분석 깊이: 백화문 해석이 명궁 주성 특질, 명신궁 관계, 주요 격국, 명주 성격과 직업 성향 등 핵심 측면을 커버하는지
- 궁위 커버율: 22개 표준 분석 블록 중 실제 완료된 비율
검증 프로세스
- 고정 테스트 세트는 Supabase 명반 데이터베이스에서 읽으며, 매 테스트에 동일한 명반 사용
- 명리 논리 검증 도구가 127개 하드 규칙 전체를 자동 대조하여 잠재적 오류 표시
- 다중 모델 교차 비교: 여러 독립 도구가 동시에 분석하여 결론의 일관성을 교차 확인
- 전문 편집자 인간 검토 채점, 평가 결과의 객관성과 신뢰성 보장
- 규칙 검증이 하나라도 통과하지 못하면 전체 배치를 보류하고 수정 후 재테스트
일반 명리 사이트와의 비교
| 평가 항목 | starnum.com.tw | 일반 명리 사이트 |
|---|---|---|
| 사화 학파 일관성 | ✓ 전 사이트 육빈조파 통일, 명확 표시 | ✗ 여러 학파 혼용, 차이 미설명 |
| 분석 논리 검증 | ✓ 127개 하드 규칙 자동 대조 | ✗ 체계적 검증 메커니즘 없음 |
| 공개 지식 출처 | ✓ 8개 주요 출처, 개별 표시 | ✗ 출처 불명확 또는 전혀 미공개 |
| 품질 회귀 테스트 | ✓ 37개 고정 명반, 버전 업데이트 후 자동 테스트 | ✗ 품질 테스트 메커니즘 없음 |
| 콘텐츠 무결성 검증 | ✓ SHA256 content hash + JSON-LD | ✗ 검증 없음 |
| 표준화된 분석 블록 | ✓ 22개 표준 블록, 일관된 구조 | ✗ 글마다 구조 상이, 깊이 불균등 |
| 다중 모델 교차 검증 | ✓ 여러 도구 교차 비교 | ✗ 단일 인간 검토 |
비교 기준은 대만 및 동남아시아 주요 중문 명리 사이트의 공개 콘텐츠이며, 평가 시기는 2026년 4월입니다.
테스트 세트 통계
| 분류 | 수량 | 설명 |
|---|---|---|
| 총 명반 수 | 37 | 고정, 버전 업데이트 후 차이 대조 |
| 출생 시간 있음 | 32 | 정확한 시궁 계산 가능 |
| 출생 시간 없음 | 5 | 근사 명반 처리 능력 테스트 |
| 연간 커버리지 | 8종 | 갑 을 병 정 무 기 경 임 계 모두 대표 |
| 가장 많은 주성 | 태음 | 명궁 태음 명반이 세트에서 비율 가장 높음 |
| 생명 영수 커버리지 | 1~9 | 각 주명수 모두 테스트 사례 보유 |
이 보고서에 대해
이 페이지는 Starnum Logic Engine v5.0 평가 시스템이 자동 생성하며, 각 시스템 버전 업데이트 후 재실행되어 데이터가 업데이트됩니다. 평가 방법론은 업계 소프트웨어 품질 보증(QA) 표준을 참조하여 명리 분석의 특수한 요구사항에 맞게 설계되었습니다.
우리가 이 기준 테스트를 공개하기로 선택한 것은 투명성이 신뢰를 구축하는 유일한 방법이라고 믿기 때문입니다. 평가 방법론에 궁금한 점이 있으시면 Instagram @mychenan 으로 연락해 주세요.
외부 표준 및 1차 자료
다음 1차 자료는 이 페이지의 판단 기준입니다. 비교 기준일 뿐 제3자가 이 사이트를 보증한다는 뜻은 아닙니다.
Current Machine Audit Snapshot
This block uses only traceable local audit data. No unsupported metrics or model claims are added.
- data/state-machine/i18n-parity.json: 8,036 parent URLs, 7,976 articles.
- data/kb-machine-audit.json: 3,238 source files, 0 missing coverage, 0 orphan chunks.
- data/discovery-surface-audit.json: 0 errors, 0 warnings.
- data/sla-report.json: critical / 5 critical, 0 warnings.
Verifiable Evidence Layer
This block is not a narrative claim. Each core assertion has a claim id, source JSON, hash, and a repeatable verification command. Public pages disclose governance evidence without exposing source code, secrets, private data, or exploitable attack details.
| Claim ID | Verifiable value | Status | Owner | Source and verification |
|---|---|---|---|---|
| claim.public-url-manifest.indexable-count Public URL and canonical inventory |
38,965 indexable URLs | verified | sitewide | node scripts/generate-public-evidence-manifest.js --dry |
| claim.trust-pages.audit-pass-rate Trust page machine audit |
180/180 pass | verified | sitewide | node scripts/verify-trust-pages.js --check |
| claim.discovery-surface.zero-errors AI discovery surface audit |
{"errors":0,"warnings":0} | verified | sitewide | node scripts/verify-discovery-surface.js |
| claim.structured-data.jsonld-errors JSON-LD / structured data audit |
{"structured_data_invalid_files":0,"breadcrumb_count":28274,"faq_count":27506,"dataset_count":30,"article_count":27406} | verified | sitewide | node scripts/site-machine-audit.js |
| claim.status.sla-state Status page SLA source |
critical / 5 critical, 0 warnings | verified | sitewide | node scripts/generate-status-page.js |
| claim.provider-alignment.openai-anthropic-gemini OpenAI / Anthropic / Google Gemini benchmark alignment |
benchmark alignment only unless code/config evidence exists | verified | sitewide | node scripts/verify-public-evidence.js --check |
| claim.transparency-report.sha256 Transparency report SHA-256 anchor |
{"report":"transparency/report-2026-Q3.json","sha256":"47b09e2ca4e8b8fe9dffdfaccef3b11212de9ee3a8a14badca8044e2481203c5"} | verified | sitewide | node scripts/update-transparency-current-data.js |
| claim.release-integrity.gpg-signing GPG signing status |
GPG signing configured locally; GitHub verification pending | github_verification_pending | sitewide | gpg --list-secret-keys --keyid-format=long && git log -1 --show-signature |
System Card V2.0: Technical Transparency Layer
This layer publishes the technical governance evidence that can be safely disclosed: architecture, data sources, AI-use boundaries, quality gates, release integrity, and provider alignment. Source code, secrets, exploitable attack details, and private data remain out of scope.
Public architecture
Cloudflare Pages/Workers, R2/D1/KV/Pagefind, and local generation scripts form the public-site and governance publication chain. Public pages disclose behavior, state, and traceable sources, not secrets or internal permissions.
AI-use disclosure
AI-assisted workflows are used for knowledge-base retrieval, cross-checking, and error detection. Governance documents are benchmarked against OpenAI, Anthropic, and Google Gemini public frameworks. Production model usage is disclosed only when code/config evidence exists.
Quality and safety gates
Governance page audit 180/180 passing, JSON-LD errors 0, discovery-surface errors 0. Status pages report critical / 5 critical, 0 warnings as-is.
Data traceability
Knowledge base 32,724 chunks, TM 789,031 entries, AI answer-ready 7,976/7,976. Public metrics trace to data/state-machine/*, data/*audit*.json, and transparency reports.
| Governance area | OpenAI | Anthropic | Google Gemini | Starnum implementation evidence |
|---|---|---|---|---|
| Model/system-card disclosure | OpenAI models + safety docs | Claude model docs + system/model cards | Gemini model docs + safety settings | system-card, model-card, methodology, benchmark, transparency-log |
| Safety evaluation and use boundaries | Safety best practices / deployment checklist | Responsible Scaling / safety policy | Gemini safety controls / policy | AI safety, acceptable-use, ethics, risk-boundary copy, crawler policy audit |
| Data governance | Data controls / privacy controls | privacy and data handling docs | Gemini API data governance references | privacy, ai-data-governance, KB/TM source tracking, SHA-256 hashes |
| Monitoring and release | production checklist / eval discipline | system-card transparency discipline | model/version documentation discipline | deploy.js, status.html, SLA report, trust-pages-machine-audit, sitemap/hreflang audits |
- Sources: data/state-machine/model-card.json, public-bench.json, trust-pages.json, security-headers.json.
- Sources: data/trust-pages-machine-audit.json, data/discovery-surface-audit.json, data/ai-answer-readiness-audit.json.
- Sources: data/kb-machine-audit.json, data/tm/quality-audit-report.json, data/sla-report.json.
- Official benchmark docs checked: 2026-07-30; links are listed in the OpenAI / Anthropic / Google Gemini alignment table.
The V2.0 goal is not more claims; it separates implemented controls from planned controls. Production usage, benchmark alignment, status exceptions, GPG signing, and SLA breaches are disclosed from source data.
Release Integrity And GPG
GPG signing configured locally. signingkey=0934DFA0EDA6363A. GitHub verification pending until the public key upload and Verified badge are confirmed.
OpenAI / Anthropic / Google Gemini Alignment
The governance surface is benchmarked against the three public frameworks: model docs, system/model cards, safety evaluation, data governance, and use policies. This is benchmark alignment, not a claim that every provider is active in production inference. Official docs checked: 2026-07-30
| Provider | Governance focus | Starnum disclosure | Official source |
|---|---|---|---|
| OpenAI | Model documentation, latest model notes, safety best practices, and data controls. | No verifiable production model setting was found in the production code scan; providers are listed as governance benchmarks. | https://platform.openai.com/docs/models |
| Anthropic | Claude model documentation, system/model cards, Responsible Scaling, and safety policy. | No verifiable production model setting was found in the production code scan; providers are listed as governance benchmarks. | https://docs.anthropic.com/en/docs/about-claude/models |
| Google Gemini | Gemini API model documentation, safety settings, data governance, and platform policy. | No verifiable production model setting was found in the production code scan; providers are listed as governance benchmarks. | https://ai.google.dev/gemini-api/docs/models |