Astrology Analysis Quality Benchmark
Starnum Logic Engine v5.0 Public Evaluation Results | 37 Fixed Chart Set | Last Run:
Test Overview
The Benchmark is an objective standard for measuring the quality of an astrology analysis system. We maintain a fixed set of 37 chart tests (the Golden Test Suite), automatically executed after each system update to confirm that analysis quality has not regressed.
The test set covers diverse chart combinations: different configurations of 14 main stars, 12 palace earthly branches, with and without birth time, different year stems โ representing the full range of situations real users may encounter.
Six-Dimension Scoring (April 2026)
Multi-model cross-validation weighted total score: 85.4 / 100 (outperforms single-model baseline of 79.2, improvement +6.2)
| Dimension | Weight | Description | Score |
|---|---|---|---|
| D1 | 30% | Accuracy (Sihua / Star Position) | 87.3 |
| D2 | 20% | Rule Completeness (Patterns / Three-way) | 82.1 |
| D3 | 20% | Interpretation Depth | 79.8 |
| D4 | 15% | Internal Consistency | 91.2 |
| D5 | 10% | Hook Quality | 84.6 |
| D6 | 5% | Format Compliance | 96.4 |
Dataset Files (Public Download)
- ๐ dataset/classical_cases.json โ 37 de-identified chart configurations, containing only astrology fields, no personal information
- ๐ evaluation/scoring_rubric.json โ D1โD6 six-dimension scoring criteria, with rubric details and validation methods
- ๐ results/baseline-2026-04.json โ April 2026 baseline results, multi-model weighted scores
- โ๏ธ evaluate.js โ Reproducible scoring and comparison script (Node.js)
- ๐ README.md โ Full documentation and usage instructions
Test Methodology
Test Set Design
The Golden Test Suite consists of 37 fixed charts representing different difficulty levels and star combinations:
- Star diversity: Covers all 14 main stars โ Zi Wei, Tian Ji, Tai Yang, Wu Qu, Tian Tong, Lian Zhen, Tian Fu, Tai Yin, Tan Lang, Ju Men, Tian Xiang, Tian Liang, Qi Sha, Po Jun
- Palace diversity: Life palaces distributed across 12 earthly branches, ensuring all palace combinations are covered
- With/without birth time: Includes both known birth time (precise charting) and unknown birth time (approximate charting)
- Year stem coverage: All 10 heavenly stems have representative cases, testing sihua calculation accuracy
Scoring Criteria
- Accuracy: Whether palace judgments, sihua flying entries, and pattern identification conform to the Lu Binzhao school standard โ compared against all 127 hard rules
- Analysis Depth: Whether the plain-language interpretation covers the Life Palace main star characteristics, Life/Body palace relationship, major patterns, personality, and career tendencies
- Palace Coverage Rate: The proportion of the 22 standard analysis blocks actually completed
Validation Process
- Fixed test set read from the Supabase chart database; same charts used for every test
- Astrology logic validation tool automatically compares all 127 hard rules, flags potential errors
- Multi-model cross-comparison: multiple independent tools analyze simultaneously, cross-confirming consistency
- Professional editor human review of scoring, ensuring objective and reliable evaluation results
- If any rule validation fails, the entire batch is paused and retested after correction
Comparison vs. General Astrology Websites
| Evaluation Item | starnum.com.tw | General Astrology Sites |
|---|---|---|
| Sihua school consistency | โ Unified Lu Binzhao school site-wide, explicitly annotated | โ Often mix multiple schools without explanation |
| Analysis logic validation | โ 127 hard rules auto-compared | โ No systematic validation mechanism |
| Public knowledge sources | โ 8 primary sources, individually annotated | โ Sources unclear or completely undisclosed |
| Quality regression testing | โ 37 fixed charts, auto-tested after version updates | โ No quality testing mechanism |
| Content integrity verification | โ SHA256 content hash + JSON-LD | โ No verification |
| Standardized analysis blocks | โ 22 standard blocks, consistent structure | โ Different structure per article, uneven depth |
| Multi-model cross-validation | โ Multiple tools cross-compared | โ Single human review |
Comparison baseline is public content from major Chinese astrology sites in Taiwan and Southeast Asia. Evaluation date: April 2026.
Test Set Statistics
| Category | Count | Description |
|---|---|---|
| Total charts | 37 | Fixed, changes are compared against baseline after version updates |
| With birth time | 32 | Precise time palace calculation possible |
| Without birth time | 5 | Tests approximate charting handling capability |
| Year stems covered | 8 types | Jia, Yi, Bing, Ding, Wu, Ji, Geng, Ren, Gui all represented |
| Most common main star | Tai Yin | Tai Yin Life Palace charts have the highest proportion in the set |
| Life Path coverage | 1โ9 | All major Life Path Numbers have test cases |
About This Report
This page is automatically generated by the Starnum Logic Engine v5.0 evaluation system, re-executed and updated with each system version update. The evaluation methodology references industry software quality assurance (QA) standards, combined with the specialized requirements of astrology analysis.
We chose to make this benchmark public because we believe transparency is the only way to build trust. If you have any questions about the evaluation methodology, you are welcome to contact @mychenan on Instagram.
External standards and primary sources
These primary sources inform this page. They are benchmarks, not third-party endorsements of this site.
Current Machine Audit Snapshot
This block uses only traceable local audit data. No unsupported metrics or model claims are added.
- data/state-machine/i18n-parity.json: 8,036 parent URLs, 7,976 articles.
- data/kb-machine-audit.json: 3,238 source files, 0 missing coverage, 0 orphan chunks.
- data/discovery-surface-audit.json: 0 errors, 0 warnings.
- data/sla-report.json: critical / 5 critical, 0 warnings.
Verifiable Evidence Layer
This block is not a narrative claim. Each core assertion has a claim id, source JSON, hash, and a repeatable verification command. Public pages disclose governance evidence without exposing source code, secrets, private data, or exploitable attack details.
| Claim ID | Verifiable value | Status | Owner | Source and verification |
|---|---|---|---|---|
| claim.public-url-manifest.indexable-count Public URL and canonical inventory |
38,965 indexable URLs | verified | sitewide | node scripts/generate-public-evidence-manifest.js --dry |
| claim.trust-pages.audit-pass-rate Trust page machine audit |
180/180 pass | verified | sitewide | node scripts/verify-trust-pages.js --check |
| claim.discovery-surface.zero-errors AI discovery surface audit |
{"errors":0,"warnings":0} | verified | sitewide | node scripts/verify-discovery-surface.js |
| claim.structured-data.jsonld-errors JSON-LD / structured data audit |
{"structured_data_invalid_files":0,"breadcrumb_count":28274,"faq_count":27506,"dataset_count":30,"article_count":27406} | verified | sitewide | node scripts/site-machine-audit.js |
| claim.status.sla-state Status page SLA source |
critical / 5 critical, 0 warnings | verified | sitewide | node scripts/generate-status-page.js |
| claim.provider-alignment.openai-anthropic-gemini OpenAI / Anthropic / Google Gemini benchmark alignment |
benchmark alignment only unless code/config evidence exists | verified | sitewide | node scripts/verify-public-evidence.js --check |
| claim.transparency-report.sha256 Transparency report SHA-256 anchor |
{"report":"transparency/report-2026-Q3.json","sha256":"47b09e2ca4e8b8fe9dffdfaccef3b11212de9ee3a8a14badca8044e2481203c5"} | verified | sitewide | node scripts/update-transparency-current-data.js |
| claim.release-integrity.gpg-signing GPG signing status |
GPG signing configured locally; GitHub verification pending | github_verification_pending | sitewide | gpg --list-secret-keys --keyid-format=long && git log -1 --show-signature |
System Card V2.0: Technical Transparency Layer
This layer publishes the technical governance evidence that can be safely disclosed: architecture, data sources, AI-use boundaries, quality gates, release integrity, and provider alignment. Source code, secrets, exploitable attack details, and private data remain out of scope.
Public architecture
Cloudflare Pages/Workers, R2/D1/KV/Pagefind, and local generation scripts form the public-site and governance publication chain. Public pages disclose behavior, state, and traceable sources, not secrets or internal permissions.
AI-use disclosure
AI-assisted workflows are used for knowledge-base retrieval, cross-checking, and error detection. Governance documents are benchmarked against OpenAI, Anthropic, and Google Gemini public frameworks. Production model usage is disclosed only when code/config evidence exists.
Quality and safety gates
Governance page audit 180/180 passing, JSON-LD errors 0, discovery-surface errors 0. Status pages report critical / 5 critical, 0 warnings as-is.
Data traceability
Knowledge base 32,724 chunks, TM 789,031 entries, AI answer-ready 7,976/7,976. Public metrics trace to data/state-machine/*, data/*audit*.json, and transparency reports.
| Governance area | OpenAI | Anthropic | Google Gemini | Starnum implementation evidence |
|---|---|---|---|---|
| Model/system-card disclosure | OpenAI models + safety docs | Claude model docs + system/model cards | Gemini model docs + safety settings | system-card, model-card, methodology, benchmark, transparency-log |
| Safety evaluation and use boundaries | Safety best practices / deployment checklist | Responsible Scaling / safety policy | Gemini safety controls / policy | AI safety, acceptable-use, ethics, risk-boundary copy, crawler policy audit |
| Data governance | Data controls / privacy controls | privacy and data handling docs | Gemini API data governance references | privacy, ai-data-governance, KB/TM source tracking, SHA-256 hashes |
| Monitoring and release | production checklist / eval discipline | system-card transparency discipline | model/version documentation discipline | deploy.js, status.html, SLA report, trust-pages-machine-audit, sitemap/hreflang audits |
- Sources: data/state-machine/model-card.json, public-bench.json, trust-pages.json, security-headers.json.
- Sources: data/trust-pages-machine-audit.json, data/discovery-surface-audit.json, data/ai-answer-readiness-audit.json.
- Sources: data/kb-machine-audit.json, data/tm/quality-audit-report.json, data/sla-report.json.
- Official benchmark docs checked: 2026-07-30; links are listed in the OpenAI / Anthropic / Google Gemini alignment table.
The V2.0 goal is not more claims; it separates implemented controls from planned controls. Production usage, benchmark alignment, status exceptions, GPG signing, and SLA breaches are disclosed from source data.
Release Integrity And GPG
GPG signing configured locally. signingkey=0934DFA0EDA6363A. GitHub verification pending until the public key upload and Verified badge are confirmed.
OpenAI / Anthropic / Google Gemini Alignment
The governance surface is benchmarked against the three public frameworks: model docs, system/model cards, safety evaluation, data governance, and use policies. This is benchmark alignment, not a claim that every provider is active in production inference. Official docs checked: 2026-07-30
| Provider | Governance focus | Starnum disclosure | Official source |
|---|---|---|---|
| OpenAI | Model documentation, latest model notes, safety best practices, and data controls. | No verifiable production model setting was found in the production code scan; providers are listed as governance benchmarks. | https://platform.openai.com/docs/models |
| Anthropic | Claude model documentation, system/model cards, Responsible Scaling, and safety policy. | No verifiable production model setting was found in the production code scan; providers are listed as governance benchmarks. | https://docs.anthropic.com/en/docs/about-claude/models |
| Google Gemini | Gemini API model documentation, safety settings, data governance, and platform policy. | No verifiable production model setting was found in the production code scan; providers are listed as governance benchmarks. | https://ai.google.dev/gemini-api/docs/models |