moonshotai/kimi-k3-free

moonshotai/kimi-k3-free

openai attested

Attestation smg:attest:v1:22af0ea2b6669ebe6c9e5e09af548587ae1068e2fde6bfcce4aca674926985f1

|

Overall score

100%

generation

100%

1 task

text

100%

3 tasks

medium

100%

2 tasks

qa

100%

1 task

easy

100%

1 task

The real tests (3)

Every task this model was actually run on. Re-runnable and verifiable.

#TaskStatusScore
1Kimi smoke testsuccess100%
2Math: solve a linear equationsuccess100%
3Classification — sentimentsuccess100%

The task set (3)

Prompts + scoring are snapshotted with the proof. Re-runnable without the original workspace.

#TaskJudgePromptExpected
1Kimi smoke testnormalizedReply with exactly the single word: bananabanana
2Math: solve a linear equationexactSolve for x: 3x + 7 = 22. Give ONLY the numeric value of x.5
3Classification — sentimentexactClassify the sentiment of this review as positive, negative, or neutral: "Absolutely loved it — fast, accurate, and easy to use." Answer with one word.positive

Task-set commitment: smg:task-set:v1:56029888419428a504c89279f13f9b5ce9bd5e60a64f5b10d154923c6ec08d84

Reproduce this eval

Re-run these exact tasks with your own model + key and compare against the published scores. The task set is pinned by hash, so it can't be swapped.

Reproduce with your own key (BYOK, never stored)

3 tasks · task-set pinned

This list is the attestation payload published with the run. To verify it, re-run the same tasks with the same scoring and compare the hash above.