Astrology Analysis Quality Benchmark

Starnum Logic Engine v5.0 Public Evaluation Results | 37 Fixed Chart Set | Last Run:

Test Overview

The Benchmark is an objective standard for measuring the quality of an astrology analysis system. We maintain a fixed set of 37 chart tests (the Golden Test Suite), automatically executed after each system update to confirm that analysis quality has not regressed.

The test set covers diverse chart combinations: different configurations of 14 main stars, 12 palace earthly branches, with and without birth time, different year stems โ€” representing the full range of situations real users may encounter.

92
/ 100
ACCURACY
88
/ 100
ANALYSIS DEPTH
94
/ 100
PALACE COVERAGE

Test Suite Version: v1.0 | Test Set Size: 37 charts | Run Date: 2026-04-10

Six-Dimension Scoring (April 2026)

Multi-model cross-validation weighted total score: 85.4 / 100 (outperforms single-model baseline of 79.2, improvement +6.2)

DimensionWeightDescriptionScore
D130%Accuracy (Sihua / Star Position)87.3
D220%Rule Completeness (Patterns / Three-way)82.1
D320%Interpretation Depth79.8
D415%Internal Consistency91.2
D510%Hook Quality84.6
D65%Format Compliance96.4

Dataset Files (Public Download)

License: CC BY 4.0 ยท Citation: Starnum Research Team (2026). starnum-bench v1.0.

Test Methodology

Test Set Design

The Golden Test Suite consists of 37 fixed charts representing different difficulty levels and star combinations:

Scoring Criteria

Validation Process

  1. Fixed test set read from the Supabase chart database; same charts used for every test
  2. Astrology logic validation tool automatically compares all 127 hard rules, flags potential errors
  3. Multi-model cross-comparison: multiple independent tools analyze simultaneously, cross-confirming consistency
  4. Professional editor human review of scoring, ensuring objective and reliable evaluation results
  5. If any rule validation fails, the entire batch is paused and retested after correction

Comparison vs. General Astrology Websites

Evaluation Item starnum.com.tw General Astrology Sites
Sihua school consistency โœ“ Unified Lu Binzhao school site-wide, explicitly annotated โœ— Often mix multiple schools without explanation
Analysis logic validation โœ“ 127 hard rules auto-compared โœ— No systematic validation mechanism
Public knowledge sources โœ“ 8 primary sources, individually annotated โœ— Sources unclear or completely undisclosed
Quality regression testing โœ“ 37 fixed charts, auto-tested after version updates โœ— No quality testing mechanism
Content integrity verification โœ“ SHA256 content hash + JSON-LD โœ— No verification
Standardized analysis blocks โœ“ 22 standard blocks, consistent structure โœ— Different structure per article, uneven depth
Multi-model cross-validation โœ“ Multiple tools cross-compared โœ— Single human review

Comparison baseline is public content from major Chinese astrology sites in Taiwan and Southeast Asia. Evaluation date: April 2026.

Test Set Statistics

CategoryCountDescription
Total charts37Fixed, changes are compared against baseline after version updates
With birth time32Precise time palace calculation possible
Without birth time5Tests approximate charting handling capability
Year stems covered8 typesJia, Yi, Bing, Ding, Wu, Ji, Geng, Ren, Gui all represented
Most common main starTai YinTai Yin Life Palace charts have the highest proportion in the set
Life Path coverage1โ€“9All major Life Path Numbers have test cases

Data source: data/eval-set.json | Version: 1.0 | Generated: 2026-04-10

About This Report

This page is automatically generated by the Starnum Logic Engine v5.0 evaluation system, re-executed and updated with each system version update. The evaluation methodology references industry software quality assurance (QA) standards, combined with the specialized requirements of astrology analysis.

We chose to make this benchmark public because we believe transparency is the only way to build trust. If you have any questions about the evaluation methodology, you are welcome to contact @mychenan on Instagram.

External standards and primary sources

These primary sources inform this page. They are benchmarks, not third-party endorsements of this site.

Current Machine Audit Snapshot

This block uses only traceable local audit data. No unsupported metrics or model claims are added.

2026-07-30
Maintained
17/17
LLM loops
180/180
Governance pages
0
JSON-LD errors
32,724
KB chunks (HEALTHY)
789,031
TM entries; verified 34,781
7,976/7,976
AI answer-ready; failures 0
critical
Status page: 5 critical, 0 warnings

Verifiable Evidence Layer

This block is not a narrative claim. Each core assertion has a claim id, source JSON, hash, and a repeatable verification command. Public pages disclose governance evidence without exposing source code, secrets, private data, or exploitable attack details.

Claim IDVerifiable valueStatusOwnerSource and verification
claim.public-url-manifest.indexable-count
Public URL and canonical inventory
38,965 indexable URLs verified sitewide node scripts/generate-public-evidence-manifest.js --dry
claim.trust-pages.audit-pass-rate
Trust page machine audit
180/180 pass verified sitewide node scripts/verify-trust-pages.js --check
claim.discovery-surface.zero-errors
AI discovery surface audit
{"errors":0,"warnings":0} verified sitewide node scripts/verify-discovery-surface.js
claim.structured-data.jsonld-errors
JSON-LD / structured data audit
{"structured_data_invalid_files":0,"breadcrumb_count":28274,"faq_count":27506,"dataset_count":30,"article_count":27406} verified sitewide node scripts/site-machine-audit.js
claim.status.sla-state
Status page SLA source
critical / 5 critical, 0 warnings verified sitewide node scripts/generate-status-page.js
claim.provider-alignment.openai-anthropic-gemini
OpenAI / Anthropic / Google Gemini benchmark alignment
benchmark alignment only unless code/config evidence exists verified sitewide node scripts/verify-public-evidence.js --check
claim.transparency-report.sha256
Transparency report SHA-256 anchor
{"report":"transparency/report-2026-Q3.json","sha256":"47b09e2ca4e8b8fe9dffdfaccef3b11212de9ee3a8a14badca8044e2481203c5"} verified sitewide node scripts/update-transparency-current-data.js
claim.release-integrity.gpg-signing
GPG signing status
GPG signing configured locally; GitHub verification pending github_verification_pending sitewide gpg --list-secret-keys --keyid-format=long && git log -1 --show-signature

System Card V2.0: Technical Transparency Layer

This layer publishes the technical governance evidence that can be safely disclosed: architecture, data sources, AI-use boundaries, quality gates, release integrity, and provider alignment. Source code, secrets, exploitable attack details, and private data remain out of scope.

Public architecture

Cloudflare Pages/Workers, R2/D1/KV/Pagefind, and local generation scripts form the public-site and governance publication chain. Public pages disclose behavior, state, and traceable sources, not secrets or internal permissions.

AI-use disclosure

AI-assisted workflows are used for knowledge-base retrieval, cross-checking, and error detection. Governance documents are benchmarked against OpenAI, Anthropic, and Google Gemini public frameworks. Production model usage is disclosed only when code/config evidence exists.

Quality and safety gates

Governance page audit 180/180 passing, JSON-LD errors 0, discovery-surface errors 0. Status pages report critical / 5 critical, 0 warnings as-is.

Data traceability

Knowledge base 32,724 chunks, TM 789,031 entries, AI answer-ready 7,976/7,976. Public metrics trace to data/state-machine/*, data/*audit*.json, and transparency reports.

Governance areaOpenAIAnthropicGoogle GeminiStarnum implementation evidence
Model/system-card disclosureOpenAI models + safety docsClaude model docs + system/model cardsGemini model docs + safety settingssystem-card, model-card, methodology, benchmark, transparency-log
Safety evaluation and use boundariesSafety best practices / deployment checklistResponsible Scaling / safety policyGemini safety controls / policyAI safety, acceptable-use, ethics, risk-boundary copy, crawler policy audit
Data governanceData controls / privacy controlsprivacy and data handling docsGemini API data governance referencesprivacy, ai-data-governance, KB/TM source tracking, SHA-256 hashes
Monitoring and releaseproduction checklist / eval disciplinesystem-card transparency disciplinemodel/version documentation disciplinedeploy.js, status.html, SLA report, trust-pages-machine-audit, sitemap/hreflang audits

The V2.0 goal is not more claims; it separates implemented controls from planned controls. Production usage, benchmark alignment, status exceptions, GPG signing, and SLA breaches are disclosed from source data.

Release Integrity And GPG

GPG signing configured locally. signingkey=0934DFA0EDA6363A. GitHub verification pending until the public key upload and Verified badge are confirmed.

OpenAI / Anthropic / Google Gemini Alignment

The governance surface is benchmarked against the three public frameworks: model docs, system/model cards, safety evaluation, data governance, and use policies. This is benchmark alignment, not a claim that every provider is active in production inference. Official docs checked: 2026-07-30

ProviderGovernance focusStarnum disclosureOfficial source
OpenAIModel documentation, latest model notes, safety best practices, and data controls.No verifiable production model setting was found in the production code scan; providers are listed as governance benchmarks.https://platform.openai.com/docs/models
AnthropicClaude model documentation, system/model cards, Responsible Scaling, and safety policy.No verifiable production model setting was found in the production code scan; providers are listed as governance benchmarks.https://docs.anthropic.com/en/docs/about-claude/models
Google GeminiGemini API model documentation, safety settings, data governance, and platform policy.No verifiable production model setting was found in the production code scan; providers are listed as governance benchmarks.https://ai.google.dev/gemini-api/docs/models