Categories
Education

The IB Diploma: Show Me What You Know

I was talking to a group of Grade 12 students last week about how they were feeling about AI. More specifically, we were talking about the IB Diploma and the increasingly complicated business of deciding what students should and shouldn’t use gen-AI to do.

The question wasn’t whether using AI was cheating; we’ve played that conversation to death. No, there has been a subtle twist in the AI conversation, and they are now asking a much more awkward question:

“Why would we choose to be disadvantaged?”

I gave the sort of answer you would expect from a school principal on the back foot…

…There are legitimate ways to use AI and illegitimate ones…using it to understand something is different from asking it to do the thinking for you…productive struggle matters; learning is supposed to be difficult sometimes…if you outsource the struggle, you may also outsource the learning.

The slightly uncomfortable thing was that the students agreed with me. But that wasn’t really the point.

In just a few months, these students will sit IB examinations taken by students all over the world. The exam papers will be the same, the mark schemes will be the same, phones will be put away, and pens and pencils will be at the ready. Whatever the imperfections of examinations, students broadly understand the deal.

They are, however, less certain about everything that happens before then.

In chat rooms and through the international rumour mill, they’ve heard about “everyone” writing their Internal Assessments or Extended Essays using gen-AI. They’ve also heard that some schools are turning a blind eye to such cheating. My students hear me when I say that these are sweeping generalisations and outrageous assumptions, and that they should assume positive intent. But they often can’t get past the idea that their fidelity to academic integrity might also be materially disadvantaging them.

There is some game theory in play here. My students are describing a type of Prisoner’s Dilemma. Both students (the one who holds tightly to academic integrity and the one who does not) might prefer a world in which neither takes a questionable advantage, but uncertainty about what the other is doing plays havoc in their minds. If I think you are playing by roughly the same rules as me, there isn’t much of a problem. But if I suspect you might be gaining an advantage I have voluntarily chosen not to take, doing the right thing starts to feel rather different.

I think we, as educators, sometimes miss this when we explain the benefits of cognitive struggle to students as though they haven’t understood them. Of course they have. Their concern is that while they are nobly developing their thinking at their desks late at night, somebody else has finished the same piece of work in twenty minutes, gone to bed and got their nine hours of sleep.

We can tell them they will be better thinkers for having struggled, and I believe they will. But I can also understand why a Grade 12 student applying to university might regard that as rather convenient advice from their school principal. Whilst we might get the satisfaction of defending the principle, they carry all the risk if we are wrong.

I started by thinking this was predominantly an academic-integrity problem, and for a while I thought it was mostly about student anxiety and stress. However, the more I have thought about it, the more I think that schools are caught in much the same problem as our students.

It’s probably worth saying out loud that I don’t believe there is a single group of virtuous schools heroically defending academic integrity while everyone else turns a blind eye. And whilst there will inevitably be considerable variation in how schools protect productive struggle, that may have much less to do with principles than with the resource schools always seem to have least of…time (or the money that might provide that time).

A teacher with relatively small classes and the opportunity to meet students individually can closely follow an IA’s development. They can ask why an argument changed, look at drafts and notice when something doesn’t fit with what they know about a student. If the same teacher had twice as many students, perhaps they would have to do things differently.

I have recently written about authentication and detection. In sum, I think detection is a waste of time, and we should be shifting towards the authentication of student thinking: process rather than product. But robust authentication consumes time.

At present, it is largely completed at the local level, in schools, by teachers, while the qualification it protects is global. This means that we (or rather the IB, universities and employers) are expecting schools operating in very different circumstances to produce something with the same confidence that an IA submitted in a student’s name represents that student’s actual capability.

Students are aware of such variation. They have friends elsewhere and compare what they are allowed to do, what teachers ask for, and what gets checked and what doesn’t. Some of what they hear will be exaggerated or completely wrong, but perception is enough to affect behaviour, which means a school can find itself investing heavily in protecting cognitive struggle even as its students worry that this just makes the game harder for them.

That’s where I find myself.

The uncertainty and variation create pressure for schools too. If students and parents think (easier) variation is permitted elsewhere, it becomes tempting to move towards the most permissive interpretation that can reasonably be defended. Nobody needs to behave dishonestly for a system to drift.

I still believe deeply in extended student work. I’m happy to share that I like Internal Assessments and the Extended Essay. I like students spending weeks investigating something, discovering that their original research question wasn’t very good, and I like students changing their minds and occasionally becoming far too interested in something nobody else cares about.

I don’t think AI makes any of this less valuable, but it does change how confidently we can use the finished pieces as evidence. A beautifully constructed 4,000-word essay may represent a transformational learning experience, but it is also true that one can now be submitted with rather less understanding than the finished product suggests. In our new era, the essay alone doesn’t tell us as much about their independent capabilities as it once did. We have to be able to live with that.

As schools, we do two slightly different jobs. We help students become more capable, and then we certify the capability they have acquired. Gen-AI may be extraordinarily useful for the first whilst simultaneously making the second considerably harder.  And so we find ourselves in a time where AI may not have made the IB’s educational ambition less valuable, but it is now harder to certify.

This is why I’m becoming less interested in asking whether AI wrote something. I don’t think the detection arms race leads anywhere particularly useful. Even if detection became remarkably accurate, we’d still have to decide what exactly counts. Can we use it for feedback? For restructuring? Or perhaps (as I find myself doing too much) spending an hour or two arguing with an AI about my ideas before sitting down to write anything myself.

To be honest, I’m more interested in whether the student possesses the capability their work appears to demonstrate, and whether an assessment can help prove it.

So I am always wary when I hear that we can fix things simply by increasing the weighting of final examinations. Just put the students in a room without access to the technology, and the authentication problem becomes much simpler. I understand the attraction, but there is something historically ironic about solving this particular IB problem by retreating back to the examination hall.

I went back to Alec Peterson’s account of the origins of the IB Diploma, Schools Across Frontiers.  Peterson describes the early IB as “action research” to establish whether an international baccalaureate was feasible. For it to work, he argued, it needed a unified international curriculum and examination system, but it also needed universities in different countries willing to recognise the results for entry.

The IB did not just have to create an innovative education. It had to create one that universities trusted for assessment.

Peterson also records an important choice about the qualification they were building. They could look for the greatest degree of commonality among the existing European and North American curricula and examinations, or they could build upon the ideas of reformers already dissatisfied with those systems and use the new qualification to try something different. The European Baccalaureate had essentially taken the first route.

But “We followed the second,” Peterson writes.

The Diploma therefore wasn’t created by averaging what already existed. Its founders were prepared to experiment because they thought education could be better. Close to 60 years later, the ambition still feels remarkably appealing, and we know that the ability to question, investigate, evaluate, and think critically and independently remains as important as ever.

What has changed, however, is how confidently we can infer those capabilities from some of the evidence we traditionally used to assess them.

The Extended Essay, for example, can remain an extraordinarily valuable learning experience even while becoming less reliable, by itself, as evidence of independently attributable capability. The same can be true of an Internal Assessment. Gen-AI hasn’t necessarily reduced their learning value, but it has reduced the confidence with which the finished artefact alone can be treated as evidence of what the student can independently do.

I think that puts the IB surprisingly close to the problem its founders were solving in the first place: how do you construct assessment capable of supporting an ambitious conception of education while producing evidence that universities can trust?

The answer is not found in more examinations. But it may mean building more moments of authentication into the assessment we already value.

A mathematics exploration could be followed by questions about the reasoning, perhaps changing an assumption and seeing what happens. An essay could require a short defence of its argument. A history student might have to respond to a new source through the argument they have developed; a science student might be given altered data and asked what changes, etc. Of course, different subjects would require different approaches because the claimed capability differs.

None of this is straightforward or new. Vivas and other authenticated assessments (which the IB has adopted and retired at various junctures over those 60 years) raise their own questions about language, neurodiversity, cultural expectations, consistency among examiners, and workloads. It’s obvious that a badly designed authentication system could simply replace one form of inequity with another. But I see these as design problems that the IB may need to solve, rather than reasons to let the underlying problems percolate.

What interests me most in all of this is the change in emphasis. Instead of trying to reconstruct exactly how 4,000 words were produced, we need to establish whether the student possesses the capability those 4,000 words appear to represent.  We’ve been talking a lot about this in my school, and I’m sure others have too.

This might also create a healthier relationship with AI. If it helps a student understand something, challenge an argument, practise or discover what they don’t know, that’s great. If it genuinely makes them more capable, that capability should also survive authentication. What ultimately matters is what remains when the assistance is absent.

And this cannot simply become another expectation placed upon individual IB schools.

We don’t assume that thousands of IB teachers around the world will interpret assessment criteria in exactly the same way simply because they are IB teachers. We standardise, sample and moderate because variation is inevitable, and a global qualification requires confidence in what its marks mean.

The IB already standardises how student work is judged. With Gen-AI, it may mean it increasingly needs to standardise how confidence in student ownership is established too.

Authentication may now require the same level of attention as moderation.

However, simply asking schools to authenticate more rigorously risks making confidence in student work depend on whether they have enough time (or money, or both). A global qualification shouldn’t require schools to be equally well-resourced to provide equivalent confidence.

If the IB cannot provide that confidence, might schools eventually have an incentive to provide it themselves?

I hope not.

But I can imagine being tempted. If we invested heavily in our own authentication, why wouldn’t we tell our university partners? We know our students, we have systems in place to follow their work closely, we question them…so…

You can be very confident in our students’ results and capabilities.

A competitor school then has an incentive to make the same claim. Perhaps a market develops for a credentialled, independent authentication service? Imagine if universities started paying attention to them. Before long, a 38 is still a 38, but we have begun inviting universities to place different levels of confidence in it depending on where or how it was authenticated.

For any international qualification, that would be a serious problem. The IB should carry that trust on behalf of its schools and shouldn’t need to explain why a particular Diploma might be trusted more than somebody else’s.

Absolutely no one wins that game.

There is a market dimension to this too.

A-levels and Advanced Placement are not immune from gen-AI. Students taking them have access to exactly the same tools. However, their greater reliance on controlled external assessment does give them some insulation from this particular problem. A larger proportion of what matters is produced under exam conditions where assistance is removed, and attribution is relatively straightforward.

So there is an uncomfortable irony here. The very richness that has distinguished the IB Diploma (critical thinking, inquiry, independent research and work developed over time) is precisely what gen-AI is now making more difficult for us to certify in Diploma students.

If critical thinking and independent inquiry, for example, are part of the Diploma’s competitive advantage, then they are precisely what the IB now has to protect and authenticate.

There has to be enough confidence in that authentication for students to believe that playing by the rules doesn’t mean choosing to lose, and for schools to believe that investing in academic integrity isn’t placing their students at a disadvantage. That is difficult to achieve when the qualification is global but much of the responsibility for establishing ownership remains localised.

I think the IB’s first response to gen-AI was right. Trying to ban it would have been futile; young people need to learn to use these tools intelligently, and schools need sensible boundaries around their use. The harder question is whether the current assessment architecture can provide sufficient confidence in what the individual student can actually do.

The IB may already have created somewhere to start working on that problem.

Systems Transformation is a new 300-hour transdisciplinary DP course that can replace two standard-level subjects. It is currently being piloted in four schools, with the IB planning to widen the pilot by around 20 additional schools in 2028 before intended mainstream availability in 2030. Its assessment deliberately uses case-study, project and portfolio work, while the IB’s wider 16+ review explicitly describes the Systems Transformation pilot as innovating with “new evidencing and assessment methods.”³

That makes Systems Transformation interesting for another reason. It provides an opportunity to innovate not only in what the Diploma assesses, but also in how the IB establishes confidence that the capabilities being assessed genuinely belong to the student.

As one of the schools fortunate enough to be involved in the pilot, I can see considerable promise in that approach.

In the UWCSEA version of the pilot, students undertake collaborative case-study work followed by individual presentations. Another assessment requires an individual project proposal to be pitched to a panel for feedback. The course is assessed entirely through coursework, with externally assessed project and portfolio components alongside internally assessed components.⁴

These assessment choices were not designed primarily as AI authentication mechanisms, and it would be wrong to pretend otherwise. That said, I think there is now a blueprint for how the DP might evolve over the next few years.

Indeed, Systems Transformation makes the problem particularly interesting because it is so deliberately authentic. As students collaborate, undertake projects, work on real-world problems, and create portfolios over time, trying to exclude gen-AI from that kind of learning makes little sense to me.

Systems Transformation therefore offers the IB something unusually valuable: not necessarily an answer to the authentication problem, but somewhere to experiment with one. It can test new forms of assessment, questioning, panel interaction or sampled authentication; examine the workload they create, how students experience them, whether they work fairly across cultures and contexts, and whether they increase confidence in the resulting assessment.

If that learning can inform the wider Diploma, the pilot may ultimately matter for considerably more than one new course.

Anyway…

I keep coming back to the question those Grade 12 students asked me.

Why would we choose to be disadvantaged?

I don’t think we should ask them to.

Students should be able to use contemporary tools to learn, research, question and create. They should also be able to trust that the qualification they are working towards can distinguish between assistance and capability.

Schools should be able to protect productive struggle without wondering whether doing so places their students at a competitive disadvantage.  And universities should be able to trust that an IB result broadly means the same thing wherever it was earned.

For most of its history, the IB has standardised the assessment of student work. In the age of generative AI, it may increasingly be necessary to standardise how student capability is authenticated, too.

The challenge is no longer simply to show us what students can produce. It is to show us what they know.

It’s a different need, needing a different, and potentially more urgent response from the IB.  


Peterson, A. D. C. (1987). Schools Across Frontiers: The Story of the International Baccalaureate and the United World Colleges. Open Court.

International Baccalaureate. “The 16+ review.”

International Baccalaureate. “Systems transformation.”


Discover more from Serendipities

Subscribe to get the latest posts sent to your email.

Leave a comment

Discover more from Serendipities

Subscribe now to keep reading and get access to the full archive.

Continue reading