Korea AI Benchmark Protocol

    The public evaluation protocol for Korean language, documents, professional workflows, deployment, and cost.

    Protocol published · Results not yet published

    The site will not display synthetic scores before tasks, graders, prompts, versions, and run logs are reproducible.

    Six evaluation tracks

    01

    Korean language

    Honorifics, ambiguity, spacing, dialect, long-form consistency, and KR↔EN translation.

    02

    Korean documents

    HWP/HWPX, Hangul tables, government templates, PDF citations, and spreadsheet fidelity.

    03

    Professional workflows

    Legal research, public administration, customer service, finance, and clinical-document tasks with domain review.

    04

    Korean retrieval

    Fresh Korean web search, source quality, citation correctness, and answer traceability.

    05

    Agent execution

    Multi-step browser and office tasks in Korean interfaces, including recovery from tool errors.

    06

    Deployment economics

    KRW cost, latency from Korea, data residency, on-premise availability, context limits, and API reliability.

    Open tasks

    Prompts, source documents, expected outputs, and scoring code are versioned.

    Blind review

    Professional tasks use at least two independent Korean domain reviewers.

    Full run record

    Model version, provider, date, region, settings, latency, and KRW cost are retained.

    Contribute tasks or review