한국형 AI 벤치마크 프로토콜
한국어, 문서, 전문 업무, 배포, 비용을 평가하는 공개 벤치마크 프로토콜입니다.
Protocol published · Results not yet published
The site will not display synthetic scores before tasks, graders, prompts, versions, and run logs are reproducible.
Six evaluation tracks
Korean language
Honorifics, ambiguity, spacing, dialect, long-form consistency, and KR↔EN translation.
Korean documents
HWP/HWPX, Hangul tables, government templates, PDF citations, and spreadsheet fidelity.
Professional workflows
Legal research, public administration, customer service, finance, and clinical-document tasks with domain review.
Korean retrieval
Fresh Korean web search, source quality, citation correctness, and answer traceability.
Agent execution
Multi-step browser and office tasks in Korean interfaces, including recovery from tool errors.
Deployment economics
KRW cost, latency from Korea, data residency, on-premise availability, context limits, and API reliability.
Open tasks
Prompts, source documents, expected outputs, and scoring code are versioned.
Blind review
Professional tasks use at least two independent Korean domain reviewers.
Full run record
Model version, provider, date, region, settings, latency, and KRW cost are retained.