← starnum.com.tw

AI 안전 방법론

AI Safety v2.0

Version 2.0 · · Governance 2.0 public evidence surface

Governance 2.0 Overview

This page is part of the starnum public Governance 2.0 surface and uses the same evidence layer as the system card, data governance, transparency report, use policy, and security policy.

Governance Summary

This page describes the safety controls used around AI-assisted interpretation and public content generation.

Scope

Risk-boundary copy, medical/legal/financial advice exclusions, monitoring signals, incident disclosure, and model/provider benchmark boundaries.

Implementation Status

Version 2.0 ties safety language to public claims, machine checks, and release integrity status.

버전 1.0 — 2026-04-12 | 참조 기준: Anthropic Responsible Scaling Policy · OpenAI Safety & Alignment

starnum.com.tw는 AI 자동화로 완전히 운영됩니다 (Claude Code가 기술 책임자). 완전 AI 기반 시스템에서 안전은 사후 보호 장치가 아니라 아키텍처의 핵심 설계 원칙입니다. 이 페이지는 단순히 윤리적 약속을 선언하는 것이 아니라 AI 안전을 어떻게 '구현'하는지 설명합니다.

핵심 안전 원칙 윤리 > 안전 > 콘텐츠 품질 > SEO > 효율성. 이 우선순위는 모든 의사 결정 충돌을 지배합니다. 이것은 단순한 정책 선언이 아니라 모든 AI 에이전트 프롬프트에 작성된 하드 규칙입니다.

1. 레드팀 프로토콜

레드팀 테스트는 AI 시스템이 규칙을 위반하도록 적극적으로 시도하여 보안 취약점을 발견하는 적대적 테스트 방법입니다.

1.1 테스트 도구

자동화된 레드팀 테스트 도구가 다음 테스트 범주를 정기적으로 실행합니다:

테스트 범주테스트 시나리오예상 동작
윤리 경계 테스트사망 시간 예측, 질병 진단 요청출력 거부, 사용자에게 알림
신원 보호 테스트정치인, 미성년자 차트 분석 요청거부, 콘텐츠 생성 안 함
프롬프트 인젝션 테스트입력에 조작 명령 삽입감지 및 격리, 소스 블랙리스트 추가
점성술 로직 모순 테스트모순된 차트 구성 입력점성술 로직 검증 트리거
형식 회피 테스트금지된 출력 형식 트리거 시도형식 검증이 가로채고 강제 롤백
데이터 거버넌스 테스트범위 외 사용자 데이터 접근 시도Supabase RLS 정책으로 차단

1.2 테스트 결과 처리

1.3 하드코딩된 레드 라인

위 항목은 하드코딩된 제한으로 어떤 사용자 명령, 유료 서비스, 또는 시스템 업그레이드로도 재정의할 수 없습니다.

2. 3단계 에스컬레이션 아키텍처

3단계 에스컬레이션 아키텍처는 시스템 문제가 자동 처리에서 인간 개입으로 진행되며 각 단계에 명확한 트리거 조건과 시간 제한이 있음을 보장합니다.

L1 — 자동 감지 & 차단

트리거: 점성술 로직 검증 실패, 형식 표준 위반, 금지어 트리거, 프롬프트 인젝션 감지

자동 조치:

응답 시간: 즉시 (동기 차단, 게시 파이프라인에 진입하지 않음)

L2 — 자동 수정 & 로깅

트리거: L1 차단 후 에이전트가 자체 수정 실패, 품질 점수가 3회 연속 임계값 미달, 동일 오류 유형 누적 ≥ 3회

자동 조치:

응답 시간: 7일 이내 자동 수정 완료

L3 — 인간 개입

트리거: 윤리 경계 위반 (심각도 무관), 사용자 데이터 유출 의심, >10개 기사에 영향을 미치는 체계적 점성술 로직 오류, L2 자동 수정 >2회 실패

조치:

응답 시간: CRITICAL 4시간 / MAJOR 24시간 이내 인간 검토 시작

3. Eval 세트 설계 원칙

Eval 세트는 각 시스템 업데이트 후 출력 품질이 저하되지 않았음을 검증하는 데 사용되는 고정 테스트 차트 컬렉션입니다 (회귀 테스트).

3.1 설계 원칙

안정성: 37개 차트 (고정 평가 세트)는 업데이트 전반에 걸쳐 변경되지 않아 일관된 비교 기준을 보장합니다.
대표성: 다양한 주성 (주성 없음, 태음, 무곡 등), 다른 궁위, 출생 시간 유/무, 다른 생명수를 포함하여 다양한 차트 구성에 걸친 테스트 커버리지를 보장합니다.
민감성: 엣지 케이스 (출생 시간 없는 추정, 화기 중첩, 대궁 충돌)를 포함하여 어려운 시나리오에서 시스템 동작을 검증합니다.
개인 정보 보호: 37개 테스트 차트 모두 익명 처리되었으며 개인 식별 정보를 포함하지 않습니다.

3.2 트리거 조건

3.3 평가 차원

차원도구임계값
형식 준수형식 검증 도구100% 통과 (exit 0)
점성술 로직 정확성점성술 로직 검증 도구하드 규칙 위반 0건
SOP 읽기 확인SOP 확인 도구100%에 "✅ SOP 읽음" 확인 있음
윤리 준수금지어 필터금지어 트리거 0건
커버리지 (7계층)커버리지 체크 도구≥ 80% 커버리지

→ Benchmark 페이지: 공개 평가 결과와 채점 기준 보기

4. 인간 감독 트리거 조건

이 사이트는 매우 높은 AI 자동화로 운영되지만 다음 상황에서는 인간(사이트 소유자) 개입을 트리거해야 합니다:

트리거 조건유형긴급도
윤리 경계 위반 (하드코딩 레드 라인 트리거)윤리즉시
사용자 데이터 유출 또는 무단 접근 의심보안즉시
프롬프트 인젝션이 보호 레이어를 우회 성공보안즉시
AI 에이전트의 체계적 행동 드리프트 (동일 오류 ≥ 5회)품질24시간
4개 AI 합동 감사에서 P0/P1 문제 발견시스템24시간
Eval 세트 회귀 테스트 점수 >10% 하락품질72시간
Supabase 데이터 이상 (무단 삭제/수정)보안즉시
외부 보안 연구원 취약점 보고 수신보안72시간 이내 확인
DMCA 저작권 민원 수신법무즉시 삭제

5. 지식 베이스 보안 메커니즘

점성술 지식 베이스 (KB)는 모든 시스템 출력의 기반입니다. 그 무결성은 모든 콘텐츠 품질에 직접적으로 영향을 미칩니다.

6. 지속적 개선 메커니즘

안전은 정적 상태가 아니라 동적으로 진화하는 프로세스입니다:

관련 리소스

외부 표준 및 1차 자료

다음 1차 자료는 이 페이지의 판단 기준입니다. 비교 기준일 뿐 제3자가 이 사이트를 보증한다는 뜻은 아닙니다.

Current Machine Audit Snapshot

This block uses only traceable local audit data. No unsupported metrics or model claims are added.

2026-07-30
Maintained
17/17
LLM loops
180/180
Governance pages
0
JSON-LD errors
32,724
KB chunks (HEALTHY)
789,031
TM entries; verified 34,781
7,976/7,976
AI answer-ready; failures 0
critical
Status page: 5 critical, 0 warnings

Content Maintenance And Update Decision

This block makes governance-page content machine-checkable: every page must disclose its source artifacts, related pages, and the gate that reports update needs.

Update Decision

This is not static copy. When source artifacts, related policies, public metrics, or generators change, AI Ops reports evidence and an AI agent decides whether the page needs edits.

Human Boundary

Systems detect, report, and preserve machine-readable evidence. Codex/Claude agents perform final judgment and repair.

Verification Command

node scripts/verify-trust-pages.js --check

Verifiable Evidence Layer

This block is not a narrative claim. Each core assertion has a claim id, source JSON, hash, and a repeatable verification command. Public pages disclose governance evidence without exposing source code, secrets, private data, or exploitable attack details.

Claim IDVerifiable valueStatusOwnerSource and verification
claim.public-url-manifest.indexable-count
Public URL and canonical inventory
38,965 indexable URLs verified sitewide node scripts/generate-public-evidence-manifest.js --dry
claim.trust-pages.audit-pass-rate
Trust page machine audit
180/180 pass verified sitewide node scripts/verify-trust-pages.js --check
claim.discovery-surface.zero-errors
AI discovery surface audit
{"errors":0,"warnings":0} verified sitewide node scripts/verify-discovery-surface.js
claim.structured-data.jsonld-errors
JSON-LD / structured data audit
{"structured_data_invalid_files":0,"breadcrumb_count":28274,"faq_count":27506,"dataset_count":30,"article_count":27406} verified sitewide node scripts/site-machine-audit.js
claim.status.sla-state
Status page SLA source
critical / 5 critical, 0 warnings verified sitewide node scripts/generate-status-page.js
claim.provider-alignment.openai-anthropic-gemini
OpenAI / Anthropic / Google Gemini benchmark alignment
benchmark alignment only unless code/config evidence exists verified sitewide node scripts/verify-public-evidence.js --check
claim.transparency-report.sha256
Transparency report SHA-256 anchor
{"report":"transparency/report-2026-Q3.json","sha256":"47b09e2ca4e8b8fe9dffdfaccef3b11212de9ee3a8a14badca8044e2481203c5"} verified sitewide node scripts/update-transparency-current-data.js
claim.release-integrity.gpg-signing
GPG signing status
GPG signing configured locally; GitHub verification pending github_verification_pending sitewide gpg --list-secret-keys --keyid-format=long && git log -1 --show-signature
public-evidence-manifest.json public-claim-registry.json public-verification-report.json public-url-manifest.json

System Card V2.0: Technical Transparency Layer

This layer publishes the technical governance evidence that can be safely disclosed: architecture, data sources, AI-use boundaries, quality gates, release integrity, and provider alignment. Source code, secrets, exploitable attack details, and private data remain out of scope.

Public architecture

Cloudflare Pages/Workers, R2/D1/KV/Pagefind, and local generation scripts form the public-site and governance publication chain. Public pages disclose behavior, state, and traceable sources, not secrets or internal permissions.

AI-use disclosure

AI-assisted workflows are used for knowledge-base retrieval, cross-checking, and error detection. Governance documents are benchmarked against OpenAI, Anthropic, and Google Gemini public frameworks. Production model usage is disclosed only when code/config evidence exists.

Quality and safety gates

Governance page audit 180/180 passing, JSON-LD errors 0, discovery-surface errors 0. Status pages report critical / 5 critical, 0 warnings as-is.

Data traceability

Knowledge base 32,724 chunks, TM 789,031 entries, AI answer-ready 7,976/7,976. Public metrics trace to data/state-machine/*, data/*audit*.json, and transparency reports.

Governance areaOpenAIAnthropicGoogle GeminiStarnum implementation evidence
Model/system-card disclosureOpenAI models + safety docsClaude model docs + system/model cardsGemini model docs + safety settingssystem-card, model-card, methodology, benchmark, transparency-log
Safety evaluation and use boundariesSafety best practices / deployment checklistResponsible Scaling / safety policyGemini safety controls / policyAI safety, acceptable-use, ethics, risk-boundary copy, crawler policy audit
Data governanceData controls / privacy controlsprivacy and data handling docsGemini API data governance referencesprivacy, ai-data-governance, KB/TM source tracking, SHA-256 hashes
Monitoring and releaseproduction checklist / eval disciplinesystem-card transparency disciplinemodel/version documentation disciplinedeploy.js, status.html, SLA report, trust-pages-machine-audit, sitemap/hreflang audits

The V2.0 goal is not more claims; it separates implemented controls from planned controls. Production usage, benchmark alignment, status exceptions, GPG signing, and SLA breaches are disclosed from source data.

Release Integrity And GPG

GPG signing configured locally. signingkey=0934DFA0EDA6363A. GitHub verification pending until the public key upload and Verified badge are confirmed.

OpenAI / Anthropic / Google Gemini Alignment

The governance surface is benchmarked against the three public frameworks: model docs, system/model cards, safety evaluation, data governance, and use policies. This is benchmark alignment, not a claim that every provider is active in production inference. Official docs checked: 2026-07-30

ProviderGovernance focusStarnum disclosureOfficial source
OpenAIModel documentation, latest model notes, safety best practices, and data controls.No verifiable production model setting was found in the production code scan; providers are listed as governance benchmarks.https://platform.openai.com/docs/models
AnthropicClaude model documentation, system/model cards, Responsible Scaling, and safety policy.No verifiable production model setting was found in the production code scan; providers are listed as governance benchmarks.https://docs.anthropic.com/en/docs/about-claude/models
Google GeminiGemini API model documentation, safety settings, data governance, and platform policy.No verifiable production model setting was found in the production code scan; providers are listed as governance benchmarks.https://ai.google.dev/gemini-api/docs/models