Chính sách Quản trị Dữ liệu AI
Data Governance v2.0Version 2.0 · · Governance 2.0 public evidence surface
Governance 2.0 Overview
This page is part of the starnum public Governance 2.0 surface and uses the same evidence layer as the system card, data governance, transparency report, use policy, and security policy.
Governance Summary
This page explains how data sources, knowledge-base chunks, translation memory, and public artifacts are tracked.
Scope
Knowledge-base lineage, translation memory audit counts, privacy-related processors, source hashes, and public evidence manifests.
Implementation Status
Version 2.0 connects data governance content to machine-readable manifests and verification commands.
Vòng đời dữ liệu hiện tại (lấy mã triển khai làm chuẩn)
Đây là quy tắc hiện hành có thể kiểm chứng. Nếu nội dung cũ mâu thuẫn, phần này và mã đang vận hành được ưu tiên.
- Dữ liệu lưu trữ gồm ngày sinh, giới tính, giờ và nơi sinh tùy chọn, cấu trúc lá số, kết quả luận giải, ngôn ngữ, trạng thái gói và mã truy cập/gỡ bỏ. Dữ liệu sinh là dữ liệu cá nhân và không được mô tả là đã ẩn danh.
- Bản ghi chính nằm trong Cloudflare R2, còn chỉ mục và trạng thái nằm trong D1. KV là bản sao phục hồi có thời hạn. Supabase là đường dự phòng sự cố cũ, mặc định tắt, không phải đường chính hằng ngày.
- Tra cứu phía trước có hiệu lực sáu tháng với gói miễn phí, một năm với gói thường và không giới hạn với gói nâng cao hoặc đặc biệt. Khi miễn phí/thường hết hạn, nội dung lá số R2 bị xóa; chỉ giữ chỉ mục người dùng tối giản và hồ sơ đơn hàng/kiểm toán.
- Khi người dùng gỡ bỏ, hệ thống ghi trạng thái xóa mềm và chặn ngay việc tra cứu, liệt kê hoặc ghi đè công khai. Bản ghi được bảo vệ và siêu dữ liệu kiểm toán vẫn được lưu. Đây là gỡ khỏi truy cập phía trước, không phải xóa vật lý.
- Chỉ thao tác tạo hoặc dịch bằng AI thực sự được bật mới có thể gửi cấu trúc lá số hoặc văn bản cần thiết đến nhà cung cấp đã cấu hình. Tên trong bảng đối chiếu không đồng nghĩa đang dùng trong production.
src/chart-storage.js · src/api-handler.js
Tiêu chuẩn bên ngoài và nguồn sơ cấp
Các nguồn sơ cấp này định hướng nội dung trang. Đây là chuẩn đối chiếu, không phải sự chứng thực của bên thứ ba.
Current Machine Audit Snapshot
This block uses only traceable local audit data. No unsupported metrics or model claims are added.
- data/state-machine/i18n-parity.json: 8,036 parent URLs, 7,976 articles.
- data/kb-machine-audit.json: 3,238 source files, 0 missing coverage, 0 orphan chunks.
- data/discovery-surface-audit.json: 0 errors, 0 warnings.
- data/sla-report.json: critical / 5 critical, 0 warnings.
Content Maintenance And Update Decision
This block makes governance-page content machine-checkable: every page must disclose its source artifacts, related pages, and the gate that reports update needs.
Update Decision
This is not static copy. When source artifacts, related policies, public metrics, or generators change, AI Ops reports evidence and an AI agent decides whether the page needs edits.
Human Boundary
Systems detect, report, and preserve machine-readable evidence. Codex/Claude agents perform final judgment and repair.
Verification Command
node scripts/verify-trust-pages.js --check
data/kb-machine-audit.jsondata/tm/quality-audit-report.jsondata/public-evidence-manifest.json- Related governance pages: Privacy Policy · System Card · Transparency Report · Security Policy
- Update flow:
npm run update:trust-pages→npm run test:trust
Verifiable Evidence Layer
This block is not a narrative claim. Each core assertion has a claim id, source JSON, hash, and a repeatable verification command. Public pages disclose governance evidence without exposing source code, secrets, private data, or exploitable attack details.
| Claim ID | Verifiable value | Status | Owner | Source and verification |
|---|---|---|---|---|
| claim.public-url-manifest.indexable-count Public URL and canonical inventory |
38,965 indexable URLs | verified | sitewide | node scripts/generate-public-evidence-manifest.js --dry |
| claim.trust-pages.audit-pass-rate Trust page machine audit |
180/180 pass | verified | sitewide | node scripts/verify-trust-pages.js --check |
| claim.discovery-surface.zero-errors AI discovery surface audit |
{"errors":0,"warnings":0} | verified | sitewide | node scripts/verify-discovery-surface.js |
| claim.structured-data.jsonld-errors JSON-LD / structured data audit |
{"structured_data_invalid_files":0,"breadcrumb_count":28274,"faq_count":27506,"dataset_count":30,"article_count":27406} | verified | sitewide | node scripts/site-machine-audit.js |
| claim.status.sla-state Status page SLA source |
critical / 5 critical, 0 warnings | verified | sitewide | node scripts/generate-status-page.js |
| claim.provider-alignment.openai-anthropic-gemini OpenAI / Anthropic / Google Gemini benchmark alignment |
benchmark alignment only unless code/config evidence exists | verified | sitewide | node scripts/verify-public-evidence.js --check |
| claim.transparency-report.sha256 Transparency report SHA-256 anchor |
{"report":"transparency/report-2026-Q3.json","sha256":"47b09e2ca4e8b8fe9dffdfaccef3b11212de9ee3a8a14badca8044e2481203c5"} | verified | sitewide | node scripts/update-transparency-current-data.js |
| claim.release-integrity.gpg-signing GPG signing status |
GPG signing configured locally; GitHub verification pending | github_verification_pending | sitewide | gpg --list-secret-keys --keyid-format=long && git log -1 --show-signature |
| claim.ai-data-governance.kb-lineage KB source-file and chunk lineage |
{"state":"HEALTHY","counts":{"source_files":3238,"chunks":32724,"chunked_source_files":3238,"excluded_source_files":0,"missing_chunk_coverage":0,"orphan_chunks":0}} | verified | ai-data-governance | node scripts/verify-public-evidence.js --check |
| claim.ai-data-governance.tm-lineage TM quality and flagged entries |
{"total_entries":789031,"counts":{"verified":34781,"template":0,"machine":754250,"flagged":0}} | verified | ai-data-governance | node scripts/tm-quality.js --stats |
| claim.ai-data-governance.public-artifacts Public evidence artifacts deployment inventory |
{"artifacts":["/data/public-evidence-manifest.json","/data/public-claim-registry.json","/data/public-verification-report.json"]} | verified | ai-data-governance | node scripts/verify-public-evidence.js --check |
System Card V2.0: Technical Transparency Layer
This layer publishes the technical governance evidence that can be safely disclosed: architecture, data sources, AI-use boundaries, quality gates, release integrity, and provider alignment. Source code, secrets, exploitable attack details, and private data remain out of scope.
Public architecture
Cloudflare Pages/Workers, R2/D1/KV/Pagefind, and local generation scripts form the public-site and governance publication chain. Public pages disclose behavior, state, and traceable sources, not secrets or internal permissions.
AI-use disclosure
AI-assisted workflows are used for knowledge-base retrieval, cross-checking, and error detection. Governance documents are benchmarked against OpenAI, Anthropic, and Google Gemini public frameworks. Production model usage is disclosed only when code/config evidence exists.
Quality and safety gates
Governance page audit 180/180 passing, JSON-LD errors 0, discovery-surface errors 0. Status pages report critical / 5 critical, 0 warnings as-is.
Data traceability
Knowledge base 32,724 chunks, TM 789,031 entries, AI answer-ready 7,976/7,976. Public metrics trace to data/state-machine/*, data/*audit*.json, and transparency reports.
| Governance area | OpenAI | Anthropic | Google Gemini | Starnum implementation evidence |
|---|---|---|---|---|
| Model/system-card disclosure | OpenAI models + safety docs | Claude model docs + system/model cards | Gemini model docs + safety settings | system-card, model-card, methodology, benchmark, transparency-log |
| Safety evaluation and use boundaries | Safety best practices / deployment checklist | Responsible Scaling / safety policy | Gemini safety controls / policy | AI safety, acceptable-use, ethics, risk-boundary copy, crawler policy audit |
| Data governance | Data controls / privacy controls | privacy and data handling docs | Gemini API data governance references | privacy, ai-data-governance, KB/TM source tracking, SHA-256 hashes |
| Monitoring and release | production checklist / eval discipline | system-card transparency discipline | model/version documentation discipline | deploy.js, status.html, SLA report, trust-pages-machine-audit, sitemap/hreflang audits |
- Sources: data/state-machine/model-card.json, public-bench.json, trust-pages.json, security-headers.json.
- Sources: data/trust-pages-machine-audit.json, data/discovery-surface-audit.json, data/ai-answer-readiness-audit.json.
- Sources: data/kb-machine-audit.json, data/tm/quality-audit-report.json, data/sla-report.json.
- Official benchmark docs checked: 2026-07-30; links are listed in the OpenAI / Anthropic / Google Gemini alignment table.
The V2.0 goal is not more claims; it separates implemented controls from planned controls. Production usage, benchmark alignment, status exceptions, GPG signing, and SLA breaches are disclosed from source data.
Release Integrity And GPG
GPG signing configured locally. signingkey=0934DFA0EDA6363A. GitHub verification pending until the public key upload and Verified badge are confirmed.
OpenAI / Anthropic / Google Gemini Alignment
The governance surface is benchmarked against the three public frameworks: model docs, system/model cards, safety evaluation, data governance, and use policies. This is benchmark alignment, not a claim that every provider is active in production inference. Official docs checked: 2026-07-30
| Provider | Governance focus | Starnum disclosure | Official source |
|---|---|---|---|
| OpenAI | Model documentation, latest model notes, safety best practices, and data controls. | No verifiable production model setting was found in the production code scan; providers are listed as governance benchmarks. | https://platform.openai.com/docs/models |
| Anthropic | Claude model documentation, system/model cards, Responsible Scaling, and safety policy. | No verifiable production model setting was found in the production code scan; providers are listed as governance benchmarks. | https://docs.anthropic.com/en/docs/about-claude/models |
| Google Gemini | Gemini API model documentation, safety settings, data governance, and platform policy. | No verifiable production model setting was found in the production code scan; providers are listed as governance benchmarks. | https://ai.google.dev/gemini-api/docs/models |