수능 영어 고난도 문항 · B2BCSAT English item generation · B2B

변별력은
선지에서 만듭니다.
Difficulty that comes from the options, not the passage.

Claude Opus가 오답 후보를 여러 개 만들고, 실제 수험생 선택 데이터로 학습한 critic이 수험생이 가장 많이 고를 오답 4개를 고릅니다. 지문은 원문 그대로 두고 선지만 만듭니다. Claude Opus drafts many distractor candidates for a real CSAT-style passage. A critic model trained on real student answer-choice data picks the four that test-takers are most likely to choose. The passage stays untouched; only the options are generated.

평가원 빈칸 128세트로 검증Evaluated on 128 KICE blank-fill sets 선지별 예상 선택률 제공Predicted choice rate per option
빈칸 추론 · 세트 미리보기Blank inference · set preview 예시 화면Illustrative
정답Key critic이 고른 오답Critic-selected distractor
막대: 예상 선택률Bars: predicted choice rate
문제The problem

고난도 문항은 많이 필요하고, 좋은 오답은 만들기 어렵습니다.Hard items are in demand. Good distractors are hard to write.

모의고사와 문항 공모에서 고난도 수능 영어 문항 수요는 계속 있습니다. 난도를 올리는 쉬운 방법은 지문을 어렵게 만드는 것입니다. 그러나 이 방법은 학생에게 배울 점을 남기지 않습니다.Korean test-prep companies commission a steady stream of high-difficulty CSAT English items for mock exams. The easy way to add difficulty is a harder passage, but that leaves students little to learn from.

지문을 꼬면 해설할 거리가 줄어듭니다Twisted passages leave nothing to explain

읽기 힘든 지문으로 만든 난도는 해설 시간에 "어려웠다" 외에 설명할 내용을 남기지 않습니다.Difficulty from an unreadable passage gives the explanation session little more than "it was hard".

매력적인 오답은 출제자의 시간을 씁니다Plausible distractors cost writer time

오답 4개는 모두 그럴듯해야 하고, 정답과 뜻이 겹치거나 서로 중복되면 안 됩니다.All four distractors must be tempting, yet none may overlap the key or each other in meaning.

실제 반응은 시행 뒤에야 압니다Real student behavior shows up too late

어떤 오답이 실제로 표를 모을지는 시험을 치른 뒤에야 확인할 수 있습니다.Which distractor actually draws votes is normally known only after the exam is administered.

작동 방식How it works

생성은 Claude Opus가, 선별은 학습한 critic이 맡습니다.Claude Opus generates. A trained critic selects.

강한 언어 모델은 그럴듯한 오답 후보를 많이 만듭니다. 그 가운데 수험생이 실제로 고를 오답을 고르는 일은 수험생 선택 데이터로 학습한 모델이 더 잘합니다.A frontier model is good at producing many plausible candidates. Picking the ones real students will fall for is a job for a model trained on real answer-choice data.

01

지문과 정답 입력Passage + key in

고객이 정한 지문과 정답을 받습니다. 지문은 수정하지 않습니다.Customer-chosen passage and correct answer. The passage is never edited.

02

오답 후보 생성Generate candidates

Claude Opus가 지문 하나에 오답 후보 12~16개를 씁니다.Claude Opus writes 12-16 distractor candidates per passage.

Claude Opus
03

critic 선별Critic selection

critic이 후보 조합을 채점해 4개를 고릅니다. 형식 오류, 정답과 같은 뜻, 오답끼리 중복, 길이 이탈은 규칙으로 거릅니다.The critic scores candidate combinations and picks four. Rule filters drop malformed options, key paraphrases, near-duplicates and length outliers.

Laya 기반 critic (ModernBERT 421M)Laya-based critic (ModernBERT 421M)
04

5지선다 세트 출력Five-option set out

정답과 오답 4개, 선지별 예상 선택률, 오답마다 매력 근거를 함께 드립니다. 해설 작성에 바로 씁니다.Key plus four distractors, a predicted choice rate per option, and why each distractor tempts. Ready material for the explanation.

현재 지원 유형: 빈칸 추론. critic은 내부 평가를 통과했고, 생성·선별 파이프라인은 파일럿 단계입니다.Currently supported item type: blank inference. The critic has passed internal evaluation; the full generate-and-select pipeline is at pilot stage.

평가 결과Measured results

학습에 쓰지 않은 연도에서도 수험생이 고른 오답을 맞혔습니다.On held-out years, the critic finds the distractor students actually chose.

질문: 오답 4개 가운데 수험생이 가장 많이 고른 오답을 critic이 1순위로 꼽는가?Question: among four distractors, does the critic rank first the one most students actually picked?

내부 held-out 평가Internal held-out evaluation
0.711
가장 매력적인 오답 적중률Top-distractor accuracy
95% CI 0.641-0.781
0.602
학습 전 같은 모델Same model, untrained
zero-shot · 95% CI 0.516-0.688
0.408
무작위 선택 기댓값Random baseline
오답 4개 중 무작위Uniform over 4 distractors
0.446
오답 순서 일치도Distractor ranking
Kendall τ-b · zero-shot 0.293

가장 매력적인 오답 적중률Top-distractor accuracy

평가원 빈칸 128세트, 연도별 교차검증128 KICE blank sets, leave-one-year-out CV

학습한 criticTrained critic0.711
Laya zero-shotLaya zero-shot0.602
최고 단순 규칙Best simple rule0.562
규칙 6개 결합6 rules combined0.547
무작위Random0.408
· 데이터: 평가원 수능·6월·9월 모의평가 빈칸 문항 128세트(2017~2027학년도). · 방법: 한 학년도를 빼고 나머지 연도로 학습하는 연도별 교차검증 11 fold, 시드 3개 평균, 95% bootstrap 신뢰구간. · 기준 라벨: 메가스터디 공개 선지별 선택 비율. 이용자 표본 추정치이며 평가원 공식 통계가 아닙니다. · 학습한 critic은 모든 기준선보다 두 지표에서 높았고, 차이의 신뢰구간 하한이 0보다 큽니다. · Data: 128 blank-fill sets from KICE CSAT and June/September mock exams, 2017-2027 exam years. · Method: 11-fold leave-one-year-out cross-validation, mean of 3 seeds, 95% bootstrap CIs. · Labels: Megastudy's public per-option choice rates (a user-sample estimate, not official KICE statistics). · The critic beats every baseline on both metrics, with paired-difference CI lower bounds above zero.
제공 형태Offering

문항을 만드는 팀을 위한 두 가지 방식Two ways to work with us

고난도 모의고사와 문항 공모를 준비하는 교육 기업, 출제팀과 함께합니다. 최종 검수와 채택은 고객 출제팀이 합니다.Built for test-prep companies and item-writing teams preparing high-difficulty mock exams. Your editors keep final review and sign-off.

문항 세트 공급Item-set supply

지정한 지문 또는 지문 조건에 맞춰 빈칸 5지선다 세트를 납품합니다.We deliver five-option blank-inference sets for passages you provide or specify.

  • 정답 + 오답 4개 + 선지별 예상 선택률Key + 4 distractors + predicted choice rate per option
  • 오답마다 매력 근거: 해설 작성용 재료Why each distractor tempts: material for explanations
  • 후보 전체와 critic 점수 공개, 교체 요청 가능Full candidate pool with critic scores, swap on request

대상: 모의고사·문항 공모 출제팀For: mock-exam and item-commission writing teams

API 파일럿API pilot

지문과 정답을 보내면 후보 생성과 선별 결과를 JSON으로 돌려드립니다.Send a passage and key; get generated candidates and the selected set back as JSON.

  • 소량 파일럿부터 시작Start with a small pilot batch
  • 사내 출제 도구에 연결 가능한 형식A format that plugs into in-house authoring tools
  • 시행 결과를 받으면 critic 재학습에 반영Your post-exam results can feed critic retraining

대상: 자체 문항 은행과 출제 도구를 운영하는 교육 기업For: companies running their own item bank and authoring tools

함께 파일럿을 시작할 팀을 찾습니다.We are looking for pilot partners.

고난도 수능 영어 문항을 만드는 팀이라면 지문 몇 개로 결과를 먼저 확인해 보세요.If your team writes high-difficulty CSAT English items, try it on a few of your own passages first.