---
title: Programming Education After ChatGPT: Could an Old Exam Idea Be Part of the Solution?
url: https://codereps.ai/blog/programming-education-after-chatgpt-old-exam-idea
published: 2026-10-01T19:35:14Z
author: Nursultan Kabylkas
tags: [assessment, llms, cs1, random-sampled-exams, retrieval-practice, teaching]
cover: https://codereps.ai/blog/images/59dd0281913f10b7d6ba7227e9cbeb5f.png
summary: LLMs can do intro programming homework, and assigning less only means students practice less. An old Soviet exam format, the random "ticket", suggests another way: publish the whole problem pool, then randomly sample from it in proctored assessments.
---

# Programming Education After ChatGPT: Could an Old Exam Idea Be Part of the Solution?

When ChatGPT came out, something became very clear to me as a programming instructor: the way we teach introductory programming was starting to break.

A typical introductory programming course still looks roughly the same as it did twenty years ago. Students attend lectures, complete weekly labs and homework assignments, perhaps build a final project, and take a midterm or final exam. The purpose of all of this is straightforward: by the end of the course, students should be able to program.

Historically, this model worked reasonably well because homework and labs served two purposes at the same time. They allowed instructors to evaluate students, but more importantly, they forced students to practice. Even before LLMs, however, I was never completely satisfied with the amount of practice students received. If grading and instructor time were not constraints, I would want every student in an introductory programming course to solve hundreds of problems rather than dozens. Programming is a skill, and like most skills, fluency develops through repetition.

Then LLMs changed the economics of homework completely.

Today, almost any introductory programming assignment can be completed with a prompt. Traditional plagiarism detection also becomes much less useful because an LLM can generate a different implementation for every student. But I think cheating is actually the less interesting part of the problem. There is a much more subtle issue, including for motivated students who genuinely want to learn.

The natural workflow is increasingly becoming: generate first, understand later.

For an experienced engineer, this can be incredibly productive. For someone learning to program, however, it can remove exactly the experience that develops programming ability. Reading working code is not the same skill as producing working code from a blank editor. Once you have seen the solution, you cannot completely unsee it. The solution narrows the search space, suggests the structure, and removes many of the mistakes and dead ends you would otherwise have encountered.

Those mistakes are not wasted time. They are part of the learning process. You misunderstand the problem, write something that does not work, stare at it for ten minutes, print some variables, realize your mental model was wrong, restructure the solution, and finally get it working. That process gradually creates intuition. When the solution appears instantly, the assignment may be completed, but much of that learning loop has been bypassed.

## The wrong response may be to assign less

One response from instructors has been to reduce the importance of take-home assignments because we can no longer be sure who actually completed them. I understand the reasoning, but I think it risks creating an even bigger problem: students simply practice less.

We can tell students that practice is important. We can provide optional exercises, recommend LeetCode, or upload additional problem sets. But in reality, optional practice tends to remain optional. Students have several courses, deadlines, jobs, internships, and other responsibilities competing for their time. Assessment creates a very powerful incentive about where that time gets spent.

This creates an interesting design problem. We need students to practice substantially more, not less. We also need assessment to remain meaningful in an environment where AI is always available. And whatever system we create has to work not only for a class of 20 students, but potentially for 200 or 2,000.

I started thinking about this problem through an examination system that is very familiar to people from the former Soviet Union.

## The old “ticket” exam

If you grew up in a post-Soviet country, particularly if you are a little older, you are probably familiar with the word *bilet*, the examination “ticket.”

Traditionally, an instructor prepared a large collection of tickets containing questions from the course. Students exactly knew the questions or topics in advance. During the exam, a student randomly selected a ticket, received some time to prepare, and then answered the questions orally or in writing.

There were obvious weaknesses in this system. Oral examinations could be subjective. Charisma could influence the interaction with the examiner. If the pool was small, luck could matter too much. Two students could also receive questions of somewhat different difficulty.

But there is one property of the system that I find extremely attractive: the expectations are completely aligned.

The student knows exactly what they are expected to learn. The instructor knows exactly what they are going to evaluate. There is very little energy spent guessing what might appear on the exam. If there are 100 possible tickets and the draw is truly random, the rational strategy is not to predict the professor. It is to prepare all 100.

Initially, I thought about this mostly as an old Soviet examination idea. But there is actually modern educational work using a strikingly similar approach.

In 2024, Kaili Vesik and Kathleen Currie Hall at the University of British Columbia published a paper titled [*Improved student learning through active retrieval practice and random-sampled exams*](https://doi.org/10.1017/cnj.2024.24). Their implementation was in linguistics rather than programming, but the basic architecture is remarkably close to what I have been thinking about. After class sessions, students were given open-ended questions related to the material. Those known questions then formed the basis of individualized exams created by randomly sampling from the pool while controlling factors such as topic and difficulty.

What I found particularly interesting was not simply the examination mechanism, but what happened to student behavior. The authors describe students coming to office hours and tutorials after having attempted specific questions and asking for feedback. They observed fewer questions about what would appear on the exam and more questions showing that students were actually working through the material and revisiting previous problems. Their interpretation was that making the possible questions explicit encouraged effortful retrieval practice and gave students a clearer structure for studying.

That sounds remarkably similar to what I observed when experimenting with this approach in programming courses.

The evidence should not be overstated. Vesik and Hall explicitly note that this was not a controlled experiment, and their comparison of final course grades did not show a statistically significant difference between traditional and randomized examination formats. So this paper is not proof that random-sampled exams automatically improve grades. What it does provide is an interesting precedent for the idea that making the question pool visible and sampling randomly from it can structure student practice without simply “giving away the exam.”

And programming may be an unusually good domain in which to take this idea much further.

## What if we gave students 300 programming “tickets”?

Imagine an introductory programming course where students are given access to a pool of 300 carefully designed programming problems over the semester.

They are not sample questions. They are not vaguely similar to the exam questions. They are the actual pool from which assessments will be generated.

You might release them gradually: perhaps 20 problems during a particular week. Students have several days to solve them. They can attend office hours, discuss ideas with classmates, read documentation, search online, and even use ChatGPT while practicing. At the end of the week, they take a short proctored quiz where perhaps four problems are randomly selected from those 20.

Then the pool grows. The final examination samples problems from the larger set accumulated throughout the semester.

At first, giving students the possible exam questions sounds like making the course easier. I think it actually changes what the course is measuring.

In a traditional course, part of exam preparation is an exercise in prediction. Students ask which topics are important, inspect previous exams, identify what the instructor emphasized in lectures, and try to determine what kinds of questions are likely to appear. But predicting the professor is not the skill we are trying to teach.

We are trying to teach programming.

With a sufficiently large problem pool, the expectation becomes very simple: be able to solve these problems.

There is no hidden curriculum. There is no “gotcha” question. There is no need to ask, “Professor, what should we expect on the exam?” The answer is already available.

Instead, the questions become much more useful: “I solved problems 1 through 12, but I cannot figure out problem 13. Can we go through my approach?”

That is the conversation I want students to have with instructors.

## Randomness is what creates the incentive

Suppose a student receives 30 possible problems and knows that four will be randomly selected for an assessment. The student could prepare only ten and hope for the best, but now they are gambling. The safest strategy is simply to learn how to solve all 30.

The important part is that practice and assessment become aligned. The activity that produces the grade is the same activity that develops the skill: solving programming problems.

This is different from the current situation where a student may use AI to produce a perfect homework submission and then discover during a proctored final that they cannot write a basic program independently. The homework grade and the underlying ability have become disconnected.

A random-sampled system reconnects them.

## AI may actually make this possible

There is an obvious reason we have not traditionally given students 300 or 500 carefully designed problems per course: creating that many good problems is expensive.

Someone needs to write the problem statements, create reference solutions, develop unit tests, check edge cases, estimate difficulty, classify concepts, identify ambiguous wording, and maintain the collection. Doing this manually does not scale.

This is where I think generative AI becomes interesting.

Instead of using AI primarily to generate answers for students, instructors can use AI to generate orders of magnitude more opportunities for students to practice. An instructor can start with a concept such as nested loops, arrays, recursion, or dictionaries and generate families of related problems with different contexts, constraints, difficulty levels, and combinations of previously learned concepts.

The instructor's role does not disappear. If anything, curation becomes more important. AI can generate quantity, but the problems still need validation, good tests, difficulty calibration, and thoughtful sequencing. The difference is that producing a pool of hundreds of exercises becomes realistic.

We can use AI against the problem that AI itself helped create.

## What happens when students use ChatGPT during practice?

They probably will. I do not think preventing that should necessarily be the goal.

A student might ask ChatGPT for a hint. They might ask it to explain an error. They might ask for a conceptual explanation. They might even ask for the entire solution.

But there is now an important constraint: eventually, that student knows they may encounter this problem, or another randomly selected problem from the same pool, during a proctored assessment without an AI assistant.

Copying the answer therefore does not completely solve their problem. They still have to make sure they can reproduce the reasoning themselves.

The incentive changes from “How do I submit this assignment?” to “How do I make sure I can solve this myself when it is randomly selected?”

I think that is a much healthier relationship between AI and learning.

## I tried a primitive version of this

I experimented with a version of this approach in an introductory programming course before today's generation of LLMs became as capable as they are now. The implementation was much more manual, so the problem pools were not nearly as large as I would have liked.

I did not run a controlled study, so I do not want to make quantitative claims from the experience. But subjectively, the difference in student behavior was noticeable. I saw more office-hour visits and more targeted questions. Instead of students asking vague questions about what would be on the examination, they increasingly came with a specific problem they had already attempted and wanted help solving.

This is one reason I was so interested to discover the Vesik and Hall paper. Their observations in a completely different discipline sound strikingly similar: students knew the possible questions, worked through them, and came to instructors with more specific questions about the actual material.

To me, that behavioral shift may be more important than whether the immediate average course grade changes. If the system causes students to spend more time actively retrieving knowledge, solving problems, making mistakes, and asking targeted questions, then we have changed what students actually do during the semester.

And what students repeatedly do is ultimately what they become good at.

## Programming also fixes some weaknesses of the old ticket system

The Soviet-style ticket examination had several legitimate problems, but programming gives us tools to remove many of them.

Grading does not have to depend on an oral conversation with an instructor; solutions can be evaluated automatically with unit tests. Difficulty does not have to depend on luck; an assessment generator can deliberately sample one easy problem, two medium problems, and one difficult problem, or balance questions across specific concepts. A sufficiently large pool reduces the value of memorizing a tiny number of answers, while repeated practice data can help identify questions that are unexpectedly difficult or ambiguous.

The system could also adapt over time. If thousands of students repeatedly fail a particular problem because the wording is confusing, that problem can be flagged. If another problem turns out to be trivial, its difficulty estimate can change. The problem pool itself can improve as more students interact with it.

The old ticket system provided the incentive structure. Modern software can remove many of its operational weaknesses.

## This is what we are building at CodeReps.ai

The idea behind [CodeReps.ai](https://codereps.ai) grew out of this problem.

We want instructors to be able to create and manage very large pools of programming problems, give students a structured environment in which to practice them, and then run randomized assessments from the same known pool. AI can help generate and diversify exercises; automated tests can evaluate solutions; analytics can help instructors understand where students struggle; and proctored assessments can measure what students can actually do independently.

The underlying idea is deliberately simple: give students many more opportunities to write code, make the expectations completely transparent, and make actual ability the safest path to a good grade.

There are still many questions to answer. How large does the pool need to be? How much variation between problems is desirable? Should students receive identical questions during practice and assessment, or parameterized variations? How does the approach affect long-term retention? Does it work differently for beginners and advanced students? How much AI assistance should be encouraged during practice? These are questions that should be tested rather than answered by intuition alone.

But I think the broader problem is becoming increasingly difficult to ignore. AI assistants are becoming a permanent part of software development, and students absolutely need to learn how to use them. At the same time, before AI can amplify someone's engineering ability, there still needs to be ability to amplify. Students need mental models, debugging intuition, familiarity with control flow and data structures, and experience decomposing problems into programs.

Most importantly, they still need the experience of staring at a blank editor and figuring out what to write.

The answer may not be to ban AI from programming education. Instead, we may need to redesign the incentive structure around it.

An old examination idea, modern research on retrieval practice, automated assessment, and generative AI may point toward an interesting model: make expectations transparent, give students far more problems than traditional courses can provide, let them practice extensively, and then randomly ask them to demonstrate what they can actually do.

More reps. Less guessing.

That is the experiment we want to run with [CodeReps.ai](https://codereps.ai).

If you teach programming and are interested in testing this methodology in a real course, we would love to hear from you.
