
Loading summary
A
Welcome to the Report Card with Nat Malkus, the Education Policy podcast from the American Enterprise Institute. Recently, Harvard faculty voted to push back on grade inflation at the institution by capping the number of A's given to students as 20%. But according to today's guest, Harvard's new policy and grade caps in general are not the right solution. Indeed, Scott Duke Commoners argues to create better incentives for students and faculty, we need to change the current grading system itself. Scott Duke Commoners is the Seraphim Rock professor of Business Administration in the Entrepreneurial Management Unit at Harvard Business School, a faculty affiliate of the Harvard Department of Economics and the Harvard center of Mathematical Sciences and applications, and an A16Z crypto research partner along with Joshua S. Ganz. He's also the author of the recent working paper, what Does a Grade Mean? Informativeness and Strategic Manipulation of Grading Systems. Scott Duke Commoners, welcome to the Report Card.
B
Thank you so much for having me.
A
So, Scott, Harvard faculty recently voted to adopt a policy that would push back on grade inflation, and you've argued against this policy for reasons that we will get into shortly. But first, let me ask you a more basic question. Why, in your view is grade inflation a problem?
B
It's a great question and maybe we should start by thinking a little bit about the term itself, right? Just what is great inflation per se, because there's even some debate and confusion about that. I think first of all, we have seen at Harvard and many universities a trend year on year of increasing grades measured by lots of different metrics, measured by average or median grades in the class, measured by number of students getting whatever the top grade is, even some real sort of extreme statistics. Harvard has a prize that's given out to the summa graduate with the most, with the highest gpa, grade point average. And I don't remember the exact numbers, but that chart, you know, we've had sort of like, you know, 50 person ties in in, you know, in some years, like recently. So there's increases in grades as a trend and then there's inflation, which one should think of as giving out either professors or instructors giving out students receiving grades that are sort of in some clear sense higher as a function of performance than they were previously. Think like inflation in the context of currency $1 doesn't buy as much as it used to after inflation. 1a inflation doesn't mean the same thing it used to mean. And these are different, right? Note, you can or rather they're linked, but they're not exactly the same concept. Right? You could have a Situation where grades just ascend monotonically, sort of just keep going up because everyone's performance is improving so much that your old performance thresholds are no longer serious. And maybe that means you change your thresholds, but it doesn't necessarily mean just because grades go up doesn't necessarily mean they've inflated. Inflation is a specific way in which grading increases.
A
So we like grade increases, but we're a little more suspicious when we're giving them out for free.
B
Exactly.
A
Okay, so, Scott, in courses that you teach, how lenient or as severe of a grader are you?
B
Oh, gosh. Well, it varies a lot by class. First of all, my students will probably tell you, or would probably tell you that I'm a very, very like, you know, I don't know, sharp and serious sort of evaluator of work in the sense that if someone gives a terrible talk and some presentation, I will tell them it's a terrible talk. And I do that because I think it's super important to hold students. In fact, I learned the HBS has this, you know, welcome, welcome to teaching sort of training thing called start. And one of the, the strongest lessons in that, which is something I've always thought but never been able to verbalize until I took this teaching training thing, is that holding students to a very high standard is incredibly important for establishing the trust relationship between the instructor and the student. Right. It's that you're holding them to a high standard, but it's not because you want to see them fail, but rather because you know that they can succeed and they can reach a higher level than you would they might have thought they could. So I hold students to extremely high standards. How I grade them depends a lot by course, like how that, how that standard translates into a grade. So first of all, at HBS in the MBA program, we actually have a forced curve. So every course within a very narrow band gives out the exact same grade distribution. And there's a little bit of leeway. You can give a slightly higher or lower number of the top grade and the bottom grade. But mostly, you know, sort of, you're, you're, you're fixed within a very narrow range. And so I grade according to the range.
A
The.
B
In, you know, in undergraduate teaching, I've graded on a much broader scale. And in my doctoral courses, historically, I've tended, you know, I've tried to set very, very difficult assignments, but take the grading part of it out of the, out of the objective, right? The goal, it's much more a. Like if you complete the Rigorous set of, you know, doctoral level work I want the students to do. Right. First of all, the students should be there because they want to learn it. Right? By the time you're taking doctoral courses, you're doing this to develop yourself. Not, you know, because there's a, you know, you know, someone at the end who's going to tell you whether, you know, sort of, you, you achieved it or not. But also I think it's very important that if you set a standard where, you know, you want students to reach really, really hard and like, try and write like, you know, serious research papers in a field that, you know, in a subf they never worked on before, Market design, which almost all the students who take my doctoral class are, there's a space for everyone to succeed at it. And so in that course, I tend to give some very aggressive feedback. Sometimes aggressive is the wrong word. Very serious feedback sometimes. But at the end of the day, like at a typical year, all the students, you know, end up with very high grades, but that's because they've achieved the objective.
A
And has your grading changed over time? I mean, do you feel like you've been fairly consistent or do you feel like you've been. How so?
B
Yeah, I mean, so first of all, in at hbs, the grading rules have not changed since I got there. In fact, I think they're the same as when I was a student. And so there, the way I think about how to calculate grades, the way I think about, like, how to evaluate different components of student performance, all of that evolves. But at the end of the day, the grade distribution is fixed, right? So, so the way in which I aggregate student performance into grades has evolved, but the, the end of the day distribution is fixed. In my doctoral teaching in particular, I've aspired over time to make it harder. And part of that is I've, I've had students, you know, sort of push me on this. Like, you know, you should push students more on this dimension or this other dimension. Right. Like, part of the feedback I collect from students is like, where do you want more? Where do you want more help? Where do you want more coaching? And sometimes they say like, oh, I wish you'd like, you know, pushed me harder on this, like, second milestone structure, you know, part of the project. And so I have built more of that in. And also, I mean, the environment and the sort of types of preparation and tools that students are coming in with have changed. And so, you know, what is a given outcome measure of performance has changed. Right. Like, and this is Something many schools and many workplaces are struggling with. You know, for example, nowadays it's possible to use generative AI models to do all sorts of components of the research exercise. And so I think a lot about how does that change how I evaluate students in a class where their objective is to write a research paper.
A
Right.
B
Like, you know, both. How do I, like, what am I trying to train the students to do? Right. Clearly I should try and train them to use these tools well as part of the exercise of the course. Course. But I also have to get them to conceptually think hard about what using them well means and like. And so the underlying evaluation, like what it is I'm trying to teach and what I want the students to take away interacts tremendously with the assignments and the tools they have available.
A
Yeah, I can imagine it's changed quite a bit given the shifts in the tools. So according to the Crimson at Harvard, in arts and humanities courses, roughly 78% of students received A's last year. In the social sciences, the figure was 62% in the sciences and School of Engineering and applied sciences, A's made up 57% and 56%, respectively. What's your take on how these grades got to this point?
B
Can you clarify? What do you mean, how they got to this point?
A
Well, that seems like a large share of A's across the undergraduate schools in Harvard. Does that seem like a process of great inflation? I mean, how do they arrive at this place to which Harvard was somewhat alarmed?
B
Well, those are different questions. So I don't, I'm not sure that I can speak much to the statistics across different divisions of the school and so forth. I mean, in order to understand what to make of that statistic of a statistic like that, one has to understand something about, you know, how many students are in the different divisions, you know, what the courses are like, you know, sort of. There's a, there's a lot of information and I just, I just don't know.
A
Right.
B
Like I, you know, and that's actually as a side note, that's part of what makes the challenge of addressing great inflation so complicated because, you know, part of a lot of people's. Including the committee at Harvard that made the recommendation for the, for the grade cap, which I know we'll talk about. A lot of people's instincts and intuitions come from a view that like, all these courses should have the same grade distribution, but that's not necessarily Right. Right. Like a, a first year, you know, a first year introductory foreign language course at Harvard probably should not have the same grade distribution as, you know, you know, an advanced physics course because we expect that a lot, you know, because we might expect that the first year, you know, language course has the property that many of the students who take it are going to be able to perform, right. Like by virtue of the fact that they, you know, ended up in a very selective undergraduate program. We sort of expect that like the 12 students in A, in a single section of introductory Spanish are probably going to be able to succeed at that. Whereas, you know, a really hard physics course with a hundred students in it might have a much broader distribution. And so, for example, again, I don't, I don't know exactly how these things are binned by department. If like, you know, the humanities numbers are counting a large number of introductory foreign language courses, we might expect more A's there mechanically. Now, that's totally separate from the other question you asked, which is how we should think about, you know, sort of what led Harvard to this point. Because we certainly have had, you know, a, a year on year trend of increasing grades and as I mentioned, you know, sort of some amount of it. That's that, you know, sort of, you know, that, that almost reaches to a, to a point where it's a little bit surreal, right. Once, you know, there's this chart in the, you know, in the grading, the committee subcommittee on grading report where they show this, you know, sort of prize for the top summa student annually. And like, you know, sort of, if you look at it, it's like how many of students were in this prize bracket per year? And it's like 1 1, 1, 1 1. Maybe there was a tie once. 2.4.0. And then it just like shoots up, right? You have this like hockey stick, you know, sort of lots and lots of 4.0s in recent years. And you know, again, like, I wasn't on the committee that did this assessment, so I can't really speak on, on behalf of them or, or the school. But, you know, my understanding is that it's come from lots of different dimensions and some of them, you know, sort of, and many of which have natural descriptions, right. One of them is there is a secular trend in student preparation, right? Like students are, you know, coming in now with, you know, a lot of them with, with more advanced courses in high school, like, you know, sort of more, you know, especially the students I teach, right. I work with a lot of our, you know, sort of very, very right tail math and computer science and economics students and like, you know, some of them are coming in with multiple years of graduate coursework in high school, which even, even five years ago wasn't as common a thing. Right. But at the same time, there's been my understanding that the committee reported that a lot of professors and teaching fellows feel pressure to give high grades. And, and you can imagine sort of an unraveling and inflation type effect, especially if students are, you know, selecting courses on the basis that they have, you know, that they expect to be able to get higher grades in them. That on the margin induces professors, right. If they, they want students in their classes because that, you know, is then use, you know, valuable for them. Right. Like, you know, teaching a large class means, you know, more resources, more students that can like lead people on the margin to nudge their grades up or even just, you know, concern about, you know. And again, I don't, I don't know how much, I don't know how much of an effect this is, but I definitely, the report spoke to it as, as, as some having a concern that, you know, there's, there's concern about, you know, sort of, you know, how do you really decide, you know, sort of if there's, you know, if students are all doing an exercise that's not very exact, right. It's not like they're all taking a, you know, you know, a multiple choice test with precise answers for everything, but they're, they're writing essays and they're all performing about the same. You know, maybe it's really hard to decide. And so again, you sort of give students the benefit of the doubt and nudge them up, especially because that sort of makes everyone feel better, right? It's like, you know, easier on you, easier on the student. And so I do think there's some amount of that unraveling type effect as well. Like once everybody else is inflating their grades, you have a incentive, at least on the margin to do the same.
A
So I didn't go to Harvard, but you did. So this is a little bit more from as a student, how much variation do you think there is in course difficulty? And I think this matters. Now obviously there's going to be some difficulty differences between introductory courses and courses towards the end, but even sort of across different courses of study that maybe the introductory organic chemistry is going to be more complicated than, you know, some comparative math courses or econ courses, as the case may be. This applies to some of the work that you've done. So I'm wondering a, what's your sense of how the, the variation in the difficulty, how large is that variation and how well do students understand that difficulty variation?
B
It's a great question. So, first of all, I mean, at Harvard, there's lots of vertical differentiation. So, for example, we don't even only have just one organic chemistry series. There are two different organic chemistry series of sort of different degrees of complexity. And, like, I realize it's a little like organic chemistry is one of the hardest classes. You know, I've never taken it, but my friends reporting to me that it was, like, incredibly, like, serious and like, so it has. And it has this very, very difficult reputation. But the idea that there could be like, a, you know, a, you know, hard organic chemistry and a super hard organic chemistry, like, you know, that's. That happens in many different departments. We have it with, like, you know, versions of math classes as well.
A
The.
B
I took Economics 1011 when I was an undergraduate, which is one of two different advanced one, two different intermediate microeconomics courses. There's something like advanced, you know, intermediate microeconomics and advanced intermediate microeconomics or something like that. So there's a lot of vertical differentiation. And then I think. And then there's certainly also horizontal differentiation, especially across courses like in our general education program or when I was in the college, in the core, you know, including. You could. You could often take some courses that were also targeted at concentrators, what we call majors at Harvard, and some that were sort of specially created for people who were not in the concentration and so were sort of assuming a much lower level of background. And then, you know, I think by nature of the different types, sort of like, you know, by nature of the different types of work in the different classes, it's possible to build a schedule that's as easy as you want or as hard as you want in terms of the structure of material. Material. Right. You can, you know, you can. You can stack your course, your. Your course program with a lot of introductory courses of different things, or you can stack your, you know, course program with a bunch of graduate courses and. And anything in between. And the school has actually historically been very, very flexible in letting students, you know, piece together their schedules very broadly. And so I think it's possible very much for students to sort into the sort of academic program they want.
A
So, Scott, recently, Harvard faculty voted to cap the number of A's that are given out per course at roughly 20%. About 70% of the faculty who voted on this measure voted in favor of it. And you have some problems with this policy, which we'll get to. But first, can you Steel man the argument for this new policy. What do you think they get right?
B
Great question. Oh, and actually one jumping off point for that. You asked to the extent that do students know which courses are easier either in material or grading? And there's certainly. There's at least one piece of very visible data for that, which is. There's a nickname for such courses for the particularly easy courses. I forget. I think it's an initialism for something, but they call them gems. Gem. I'm. I really. I think it's an initialism. I don't remember what it stands for, but it's named. There is a, there's a descriptor for classes that are perceived as being very approachable. We'll say yes, and it shows up in the course, you know, is like students like write course reviews and things and it shows up there. So at least at that end of the distribution, you know, there is some like coordinated sense of some classes being particularly approachable. And so let me, let me now make the case that we desperately need to shut that down. You know, the, the case, sort of the case as, you know, sort of steel, you know, steel. Steel band. I guess the, the case for the, for the grade cap starts with a. We have to do something, right? It is absolutely necessary that we stop grade inflation. Grades are out of control. We are hearing from, by the way, this, this is, again, this is not my opinion as stated. But you know, we are hearing from, as reported by the committee, we are hearing from graduate programs like professional schools and some employers that they have trouble interpreting Harvard grades because they're all sort of smushed together and also high. And that they're going to start not paying attention to students transcripts or that they're at least threatening to do so if we don't figure out some way to spread the distribution. Moreover, students have taken this upon themselves as such a source of pressure that they feel it is necessary to get a perfect grade point average and that if they don't, it's like a personal failure. And because this culture of high grades is so accessible, it's it, you know, it adds extra pressure to students and thus also of course, extra pressure to faculty and teaching fellows. And you know, and that's not all right. Like, you know, it affects the sorting of students into classes, right? If students are going out and searching for, you know, gems, you know, mining, mining for the gems or something like this, then that means they're not really taking the courses that are the ones that are sort of most appealing to them from the perspective of education. But rather they're, they're going and looking for the easy A, which to distorts the, you know, enrollments across departments and courses. And also just, you know, sort of is pedagogically bad, right. It like harms their education because we want them taking courses, because those are the courses they want to take. And then last, one more thing that shows up in the, in the sort of, you know, the, the argument in, in favor of doing something. You know, we have a handbook that says an A at Harvard means extraordinary distinction. And it's just hard to believe, right? Like, you know, sort of, you know, if people are treating the A as the student successfully achieved the objectives of the course, that isn't what it's supposed to measure. Like, we claim that this is reserved for extraordinary distinction. Okay, note I haven't said anything yet that has anything to do with the grade cap.
A
Right.
B
Like, this is sort of, you know, I've just made the case that we have to do something, right. If I had to defend the grade cap, which I find very hard to do because I think it's going to actually take us backwards. But if I had to defend it, I would say the following things. One, we have to do something. And this is a very straightforward and intuitive proposal. It is clear what it means, it is clear how to adjudicate it. It is clear, you know, how to tell, you know, sort of like how to implement it at some level. Right. You know, we're just changing like a number in our system. Previously you could put arbitrary, many arbitrarily many A's on a student. Now you now, or on students now you can only put, you know, 20 plus 4. Another thing that the proponents of the grade cap have pushed very hard is the idea that just recommending a change is not good enough because there's always this pressure for faculty to deviate, right. Like on the margin. If other people are giving lower grades, you'd like to be giving slightly higher ones, maybe. And so there's a need to constrain the grades at the entire, like to take away the freedom to make decisions. Right. You actually have to like impose a, you know, a top down constraint. And then I guess the, the other argument, I guess the other argument you can make in favor is that it's very legible, right. It, it clearly indicates that we are trying to do something. And, and that alone is A, is a very important step towards actually arriving at whatever we think is a solution.
A
Yeah. So as far as I Mean in part, the public relations component and the feasibility component seem to be near at hand. Right. I mean, it makes clear that we're taking a stand and it's doable and you also think it's a mistake. So let's get to that part. What's wrong with grade caps?
B
Oh my gosh, so many things. So first of all, just at some basic, you know, it's a basic level and I've written about this in research work with my co author Joshua Gans. There's a sort of grades are solving of or grades and transcripts are solving a very difficult problem. Right. They're trying to project two dimensions, at least of information into a single dimensional object. They're combining information about a student's ability and information about the course's difficulty and projecting downwards. Right. So you're projecting something about how well did the student perform in the course, which is really a mixture of how able was the student to do well at the course's material and how complicated was that material or how hard was the grading standard? When you impose a grade cap, you actually are regressive in terms of your ability to reveal this information. So, you know, first of all, unidimensional grades don't do such a good job at it because, you know, you don't know. You know, if you see a student having gotten an A and like this is before you have professors inflating or anything, you just don't know did that student do really well in a, you know, in a hard class or did they do moderately well in a super easy class? But that still means they got the A, right? They could have been the highest graded student in the classroom in both cases, but that could be mean a really different thing.
A
So, Scott, let me just repeat this back so that I'm making sure I understand what I heard. Essentially we have one piece of information which is the grade and it's an ordinal grade. It's not, you know, a super informative thing. But we have this and it should identify how well the student did and how difficult the course is. And because it has those two dimensions, it's already a little bit of a poverty. And now we're going to cap the grades. Am I getting sort of like the central.
B
Exactly.
A
Initial challenge? Right?
B
Yes, exactly. So we're already sort of like grades are not informative about this two dimensional structure because they're projecting down to one dimension. As you say, it's a bit of a poverty. Right. We're already sort of setting ourselves back, but now we're actually going to further distort the ability of grades to provide information. And as we'll see in a second, that has incentive consequences too in terms of that student class sorting I was describing. So, you know, the easy and intuitive example around Harvard is often, the people often quote is this course called Math 55. We have this extremely rigorous, I mentioned these, like vertical differentiations within single course. So at Harvard we have math 21, 23, 25 and 55, which are all sort of nominally the multivariable calculus, sort of like same, same math course. And you know, the joke is that the course number is roughly the number of expected hours of homework a week. The, you know, so math55 has at least historically often been taken by, you know, sort of very, very elite students who come in with a lot of like formal mathematics background and, and some years do like, you know, like the, the very, the course difficulty varies somewhat stochastically by who, you know, who teaches it. But, but it's very hard and can be extremely hard in some years. And you know, but if there are say 20 students in it, or let's, let's do 10 for, for ease of the numbers, there are 10 students in it under Harvard's 20% plus four, right? That's the number of A's available under the grade cap. You have at most two plus four, six A's available for those 10 students. Same if you have, you know, 10 undergraduates in A, in a doctoral level course, like my Market Design class, right? Like you're sort of the instructor is required to indicate some of those students as the bottom and does not give them A's. And so first of all, that's just like, you know, in a world where we were already worried that grades are losing information about course difficulty, now we've destroyed the instructor's ability to give more A grades in a course where the students are stronger on average doing harder
A
work and the students in Math 55 are well beyond all the students in Math 23. I mean, we would sort of assume because for a student who might be tempted to take the lower course, math 55, it at least sounds sort of suicidal, right? It just sounds like incredibly difficult.
B
So, so first of all, I don't want to, you know, I want to be very clear and careful on two dimensions here. First of all, lots and lots of, we have lots of really good mathematician, you know, people who end up being really excellent mathematicians coming even out of math 21. So it's not that you can't be great mathematician Coming out of wherever. Second of all, you know, math 55 is hard and very serious, but hopefully not a mental health stretch for the students who take it. But yeah, there's a huge difference in difficulty in the material. And typically this also involves a difference in preparation, right? It's like the course material in Math55 is often intended, you know, to challenge people who've come in with, with a lot of even, you know, advanced undergraduate, maybe even a little bit of graduate level math when they arrive. And, you know, and, and the assignments are, are hard. You know, same thing if you're, you know, if you're taking a graduate course, you know, sort of if in a universe where, you know, 10 students take an advanced doctoral elective, like, you know, the inference is probably, you know, and if they all succeed, right, if they all do all the work and perform at the level of a doctoral student, you know, the instructor ideally would be able to give them all the top, you know, sort of the top grade, whatever that is, right? To indicate these students all performed at the doctoral level, which is higher than the undergraduate level in this material. But when you cap the number of grades, the instructor does not. A grades, the instructor does not have the leeway to do that and worse. So the proposal has two prongs. One of them is this cap on A grades. The other is ranking the students by what's called average percentile rank. So average percentile rank means for the purposes of internal honors and thesis, not thesis prizes, but like graduation prizes and things like this, students are going to be compared to each other based on where they stacked up relative to their classmates in the classroom. So, you know, so both of these effects, right? The grade cap suddenly means you're competing against your classmates, right? Your ability to get the top grade depends on how you do relative to the classmates sitting next to you. And the average percentile rank thing stretches that out and makes it even more extreme, right? Maybe all the students in, you know, all the undergraduates taking math 55 performed at what should be an A, and they all perform at a super high level or, you know, all the students taking some doctoral course do. Or conversely, all the students taking introductory language perform at the maximum level of this class. Now we're going to force separate them, right? We have to, like, describe. We have to say that some of them stick with the 10 students. At least four of them didn't, you know, sort of rank with their peers. And moreover, the instructor has to decide, well, who's the bottom? That person is counted as a zero from the perspective of average percentile rank, even if they get an A minus, which counts to their grade point average, they're a zero for the perspective of honors. And so a student who takes exclusively advanced doctoral courses, gosh, let's say you take, you know, a whole bunch of courses where you're one of the only two undergraduates and you just happen to be the second best of the undergraduates on whatever the class metrics are. In half of those classes, you're now a zero in average percentile rank and a huge share of your schedule. And so the metrics as they're described in the proposal that was adopted, are going to say you shouldn't have honors. Good. Gosh, you're like the 0th percentile 50% of the time. And that just doesn't make any sense.
A
So you argue that grade caps are going to push students away from more difficult courses. And it pretty much follows directly from what you're saying, right? I mean, exactly. You are now thrust into this situation. If you take hard courses, your transcript is not going to reflect achievement. And if you take easier courses, you are more likely to have on your transcript a summative grave that makes you look much better. Am I reading correctly?
B
Totally. That's exactly the issue. It's another form of what we call an economics unraveling, this idea that an equilibrium unravels. As a few students realize, oh gosh, I have some chance of being the bottom of the class, by the way, it doesn't mean they actually are going to be. It's a risk aversion thing, right? Because you have to decide when you sign up for classes. You walk into the room, you see some classmate, you know, who is really good at something, and you're like, oh gosh, you know, there's no way I'm going to be in the top six of this room of 10. So then maybe you decide to take the class one level down. You're looking for a pond where you are a relatively bigger fish. Enough students drop out of the top level class. Well, now the number of A's shrink and there's a new student at the bottom. And so that student might go looking. And so you can actually imagine this cascading, you know, where students sort of gradually sort of like sort of leak out of or sort of like, you know, sort of consecutively or cascade out of the top course into the second reg course that maybe, you know, then maybe even some of them push down or it pushes students out of that class. And so there's a lot of Risk. Now, when you take an advanced class, right, you are putting your grade at risk if you stretch yourself to a place where you don't feel like you can be at the top of the class.
A
And what does this do to, and just sort of this theoretical model, what does it do to faculty's motivations to make their classes more difficult or less difficult? Does it leave it unchanged?
B
That's a good question. I mean, it's, it's hard to know. There are effects going in both directions. On the one hand, I think when you force faculty members to figure out a way of making their course grading more diagnostic, that sort of forces them to do things that will introduce metrics that maybe didn't exist before. Right. Like, you know, if you have to figure out, well, how do I rank the. Like, I, I honestly I, I don't even know precisely how I would do this in my doctoral course. I guess I'm gonna have to figure it out. But like, you know, suppose I have 10 undergrads in the class and they all turn in graduate level papers. Am I deciding, like, how do I decide, like, which of these are the most advanced doctoral level? Like, you know, I've got to figure out something to do and, and that will introduce some level of like, grading difficulty, you know, from my perspective and from the student's perspective. That didn't exist before. Right. Previously, it's like, okay, if you're performing at the level of a third year doctoral student, you're good. Like, you know, I think you have, you have achieved far and above beyond the objective of the course. Whereas now, like, I actually have to sort of figure out what additional, like, markers one uses to, to differentiate. So there's that. And that, I think pushes towards more difficulty in a certain sense in that, like, courses will be more, will be forced to be more rigorous in how they make comparisons among otherwise similar students.
A
Yeah, I mean, if you're forced to give some students not A's, then you may need to change what you are offering so that some of them will indeed not get the A Right.
B
Exactly.
A
The most integrated way to do it.
B
Absolutely. Or at least change the way you evaluate. By the way, this actually I come across every year at hbs, right. Because in HBS courses, whereas I said there was this force, we have this forced curve, there's a tremendous challenge always in figuring out, you know, sort of the students who are on the margin between the, you know, sort of the, the, the on the bubble. Right. Like sort of just above or just below what would normally be the Cutoff because there's some amount of noise in the grading process, right? Like, you know, HBS students are graded in part on participation in class. And like, you know, you as the instructor, don't always see somebody's hand or something like that. And so, you know, if a student, if two students are like at that edge by, you know, and off by, I don't know, a fraction of a point, that's noise, right? Like, there's not a, there's not a clear way to differentiate just based on the, the observed metric. And so I find myself trying to figure out, like, what else can I go and find out about these students? And we're, you know, not, not literally, I'm not like, looking up like, okay, like, does this student build a marketplace? But like, what else can I think about, about the way they interacted with the class that can help me figure out that boundary? Because again, like, you know, sort of, I like participation grading is done in real time, final papers. I typically grade like, you know, sort of anonymously and then match the grades to the students afterwards. So you do as much as you can to make the system fair, but there's some amount of noise. And then at the end of the day, you have to figure out how to account for it when you, when you're forced to break ties. So there'll certainly be a lot of that. But on the other hand, there's a lot of pressure to make things easier, right? So, for example, we just talked about this unraveling effect where there's like some pressure for students to, to seek out classes where they're the big fish, you know, the other thing they're going to seek out, of course, are courses where there are more A's to go around, right? You have a higher chance in many of these contexts of, of getting the A at the easier, larger class. And so I think, you know, with all the sort of incentives that people were thinking about as concerns for why faculty may have chosen to inflate grades, I think we're going to see, you know, potentially like, you know, sort of something of a race to the bottom. If, you know, if we think that faculty members were grading relatively easily because that was a way of attracting students and having larger courses, you know, having the relatively easier courses under, you know, in equilibrium should also still be the way to get the, you know, the most students.
A
And I ask this in part because I anticipate talking about your proposed solution and eigenvalues, not eigenvalues, eigengrades. Rather, forgive the slip in no, they are eigenvalues.
B
Eigengrades is a portmanteau, fair enough.
A
But by putting a cap on the A's, you don't necessarily. There's no mechanism by which the institution learns much about the difficulty across courses. Or I mean, is there, is there, is there any signal that is a feedback loop?
B
This is one of the things that strikes me as the strangest about the proposal, right? Like, if the real objective is to better understand and align grading in classes with course difficulty, then we absolutely should not be forcing every course to have the same grade distribution. Right. I mentioned before, it's regressive. Right. If you really want it to be the case that the students in the harder courses are, you know, treated, you know, as, you know, as having performed more significantly. Right. More, you know, more extraordinary distinction. It doesn't make any sense to say that the amount of extraordinary distinction available in, you know, in, in a, you know, doctoral course is the same as the amount of extraordinary distinction available in Act 10. Or actually even more extraordinary distinction is available in Eck 10. Right. Like, it's just like the premise that Harvard courses should all have the same grade distribution is very contrary to the view that we should be treating some courses as hard and some as easy and trying to like, get grades to be a real metric of sort of performance on some cross comparable scale.
A
Yeah, well, it seems to me that one of the things, and this is not just at Harvard, I mean, so most of the work that I do is more focused on K12. So I often think about this, and in this I have a great deal of concern that what we may want to optimize for is in this same sort of general area is that course difficulty is relatively similar. If we could optimize for minimizing the variance of difficulty across courses that are all called Algebra 1, well, our education system would probably be much better than if somehow we had a similar distribution of grades across a bunch of very different algebra.
B
Yes, exactly. I completely agree.
A
So in a minute, I want to talk to you about your proposed solution. But first I want to get some grades from you on grade it. Are you ready?
B
Let's go.
A
The design of dating apps
B
circa Win
A
as they are currently.
B
Okay, so design of dating apps, by the way, for. For people listening, this sounds like a total non sequitur. But actually the very first case in my MBA market design course called Making Markets is about the design of dating apps. I would say dating apps have been something like sort of, you know, a minus to a. As a. As A form of sort of changing the way in which the dating market works and I don't know, like a, like a B minus in establishing sort of sustainable, like, sort of, you know, sort of sustainable, well optimized and well structured marketplace businesses.
A
Frequent flyer programs.
B
This is amazing. We're going, we're apparently going through my grades on the entirety of my sort of like applied market design.
A
Well, you know, we want to get, we want to draw a little bit of your diverse experience.
B
Okay. Frequent flyer programs, again, pretty great for airlines, right? Like, sort of frequent flyer programs are a way for airlines to create different. Sorry. The, the most cynical economist view of frequent flyer programs is that they are a way for airlines to create differentiation among products that otherwise have no differentiation. Right. Like, you know, you, you become attached to one specific airline and that means that they can actually like, you know, raise your effective price. So, so I'll, I'll give them a, you know, sort of a. I'm going to be a Harvard professor. I guess I'm giving out too many A's, but I'll give them an A on the airline side and on the consumer side. I don't know. I was, I was a, you know, I appreciated frequent flyer programs sort of once upon a time, but I think I've, I've, I've gradually dulled on them. So, so we'll call them, you know, sort of, you know, you know, B plus for convenience and lounge access and like maybe a B minus for consumer protection.
A
Completing a PhD in two years. You did this, right?
B
I did that. I don't know that I'm allowed to grade my own performance, but I will tell you that I am extraordinarily proud of one of my former doctoral students and longtime friends and collaborators now, Ravi Jagadisan, who completed a PhD even faster than I did, and I would definitely give him an A triple plus. And he took my doctoral course as an undergraduate and he most certainly received an A for doing doctoral level work.
A
But would you recommend the blitz strategy or are there.
B
I think first of all, I got very lucky. There was a mathematical miracle. The first thing I tried for my dissertation actually worked. And then. And so, like, before I started grad school, I had sort of found the thing that became the core of my dissertation that summer. So like, you know, this is absolutely like a huge part of this was like math miracle. And I can't even claim we knew what we were doing because we were actually trying to prove the opposite of what we ended up showing and then discover. And like, after A month of getting nowhere. We decided to flip the problem and, and it was a miracle. I. But I will tell you that part of why I graduated early was because this incredible opportunity to visit the University of Chicago at the Becker Friedman Institute appeared. And it was one of the most extraordinary things I've ever done, especially at the time I was there working with Gary Becker and Lars Hansen and Jim Heckman. It was sort of extraordinary. And I've often said that the University of Chicago turned me from an economist into an economist. Like I really credit those couple of years with a huge part of how I think as an economist today and would never have been able to do that if I, if I hadn't been lucky enough to finish my dissertation a little speedily.
A
Sports betting in America today.
B
I'm not sure I have a strong opinion on the current state of sports betting in America, but we know empirically that it has really, really strong negative externalities and is very costly for the betters. So I'm a, you know, I'm a, I'm a low grade on the value to society and I'm not totally sure on the, on the state of the
A
market, the relevance of NFTs to the average American and.
B
Absolutely.
A
And a quick definition of NFTs for the uninitiated.
B
Non fungible tokens. These I've written a book about. So by revealed preference, these are definitely getting an A. They are blockchain based tokens that give a way of having property rights over a digital good that don't rely on the platform that created them. So if you think about like a, a book you have on, on Amazon Kindle or like a song on itunes or something like this, your copy of the asset is only meaningful within the platform, the program that they created and in fact they can change the rights and policies. There's this famous example from early in the days of Kindle when they rescinded copies of 1984 from People's Kindles. Which you know, is sort of hilarious and, and you know, and very, very meta, very Big Brother, but exactly, very ironic. But, but also there's a, there's a structural problem. You know, sort of. It's hard to imagine something like Amazon going out of business today, but it could. And if it does, all of those digital assets disappear, right? They don't have existence outside of their platform. Non fungible tokens are digital assets that exist independent of a platform they live on. Sort of like a global ledger, sort of like a digital deed. And they both are useful as property rights per Se and they're useful because they enable market transactions that would never happen otherwise. So for example, actually my mom and I just had a letter to the editor of the Washington Post about this recently. They solve the ticketing secondary market problem. NFTs make it possible to uniquely own a digital ticket in such a fashion that if I transfer it to you, you can be confident that you are now the owner and I am not, and I can't sell it to anybody else. And you can even programmatically like through software, wrap the transfer of ticket with the transfer of funds. So the ticket can only move if the payment moves at the same time. And that would make, you know, the secondary market for ticket ticketing. And it has, when NFTs have been used for tickets, massively easier and more secure.
A
And so this, this relevance, I imagine its stock will rise over time. I certainly average American's actual utility.
B
I certainly think that this will become part of an infrastructure layer that we all interact with all the time. Like everything else on the Internet that like our digital goods, right? Like right now everyone owns tons of digital goods, but they don't really own them because they're locked in these platforms. You know, over time this is such a better technology, especially for market design around digital goods, that we're going to see more and more categories of goods sort of roll into this space and you'll be using them without knowing. Just like, you know, you don't think about the type of media file your, your songs are, right? Like I guess they're MP3s. Maybe they're MP4s now with a video attached. But like you don't think, you don't call them MP3s, you just call them songs or media files. Same thing here. It's like you won't think of NF tickets as NFTs. They're just gonna be tickets that'll be like how all digital tickets work one of these days.
A
Last one. Creating puzzles.
B
Oh my gosh. Yeah, I really am the Harvard professor. I'm giving, I'm giving out way too many A's here. But the, but, but I think most of these technologies have earned it. I, so I love creating puzzles. I actually, I teach like some classes on this and I have friends who do. I like, I have friends who are like full time professional puzzle designers. Like, I work with this style of puzzle called enigmatic puzzles. These are puzzles where unlike in a crossword, say, where you get a grid and a bunch of clues and you know what you're supposed to do, you're Filling in the grid. That's the entire exercise. Enigmatic puzzles. You get some stuff that looks completely confusing. You don't necessarily know what it is, and you have to figure out, what am I looking at and what does it mean and what am I trying to do to solve it. It's a little bit like an escape room, but, like, you have to find the walls of the room while you're looking at the stuff that's in the room. And so I've been. I've been solving these for, gosh, decades. And I've been with college friends,
A
we
B
have a puzzle team that competes in the MIT Mystery Hunt and some other puzzle events. And I've been writing them for a long time. I wrote for a couple of these years a puzzle column for Bloomberg. I've done them for. For university events and stuff. And, like, it's just a very different way of thinking. But it's that. That, you know, it's this funny mixture between, like, teaching, right. And teaching. You're trying to lead somebody into a set of ideas and set up the framework for them to, like, bring themselves there. And this is that, like. But in a. In a space where using a skill set they maybe never even knew they ever had, right? Like, you know, and along the way, they learn something, right? Like, oh, gosh, those weird collections of letters. They actually have, you know, weird. They have, you know, airport abbreviations in them, and, oh, whoa, it's a map. And like, you know, those sorts of discoveries and AHAs, like, I think are very powerful for people to experience. And I take tremendous joy, both intellectually and personally, in helping people find those sorts of doorways for themselves.
A
Well, we'll leave a link in the show. Notes. I tried three solved, one totally lost.
B
Which one did you solve?
A
The music puzzle. The mixed test?
B
Yes. Playlist. Okay, great. That is a fantastic starter puzzle. I saw strongly recommend people try the playlist puzzle on. And we'll give the link.
A
It was quick. All right. Speaking of solving puzzles, Scott, how do we fix this fundamental problem with graving grading? I want you to keep this sort of at a. As a higher level, we'll link to the paper, which is, what does a grade mean? Informativeness and strategic manipulation of grading systems. But give us the contours.
B
Awesome. Okay. And also, I have to do this because we're on audio. There's a secret pun in the title, right? It's what does a grade mean? The good. Okay. So the listeners won't be able to see the smirk, but I'm glad We got a smirk.
A
I'm seeing it.
B
This is the objective of putting a subtle pun in a title. We want everyone to be like, oh, come on Scott, why did you do that?
A
Clever titles are the point.
B
So.
A
Yes, indeed.
B
But okay, so first of all, let's go back to the, the, you know, sort of the way we set up this grading problem and the paper. The first thing we do is that two dimensionality point that we talked about at the very beginning. This idea that a course grade is revealing is sort of projecting two different types of information into one unidimensional signal. And it's, you know, at a very high level. It's like, you know, the student's ability. And again, it might be ability match as a function of the course. Right. Maybe you're really, really good at, at classics and you know, and not so good at, at, you know, I don't know, at, at, at Polish. And so you do really well in your class class and not dwell in your Polish class. But those are those, you know, sort of superficially might look similar but, but it's not a statement about your absolute ability. It's, it's your match type. So student ability and course difficulty. Right. And this idea that like, you know, when you see an A grade, you know, you don't, or any grade, in fact, you don't know whether the student's performance was high because they're really high ability in the course's hard or because they're really low ability, but the course was really easy or something in the middle or some mixture thereof. And moreover, if we really aspire for grades to be calibrated evenly in performance across the school, this problem is sort of fundamental, right? Like if we're going to say a given level of performance is what determines a grade, then that same level is easier to achieve in an easy class. Right? So it's not, you know, the problem is very fundamental to the way in which grades are constructed. So first of all, there's, we propose a specific solution in the paper, but I'm actually going to start with like the, the easy first cut solution that I've been giving people whenever they've asked, you know, what, you know, Scott, if you, if you were running, if you were running the sort of decision making process, you know, what would you have recommended, you know, sort of instead of a grade cap. And I've said, you know, look, first of all, I wasn't on the committee. I didn't do all the, you know, I was not involved in all the detailed research. They did so, like, you know, I'm speaking from my perspective as someone who has seen the output of the committee's work rather than all the intermediate steps. But to me, a much more straightforward solution that actually does speak to the information problem would be to just put on every transcript not just the student's grade, but what the distribution of grades in that class was. And so why does that help? It's a really coarse. That distribution of grades is a really coarse measurement, and it's not a perfect measurement. Again, we can talk about why in a sec. But it's a really coarse measurement. Of course, difficulty, pun, not originally intended. So, you know, if you see that someone got an A in, I don't know, we'll make it up, you know, you know, in, in, in economics 6, 7. So someone got a, you know, got an a in economics 67. You don't know anything about what that means in the abstract, right? Like, okay, you look at the course title and it's called, you know, sort of, you know, Economics of Economics of the Internet. But if you see next to that A, this course had 700 students in it and 700 of them got A's. You're starting to form an inference about the difficulty of Econ 67. And conversely, if you see a distribution of grades, that said, this course had 700 students in it, and of them, two got A's. Now suddenly, like, you know, this is. This A looks very diagnostic, right? Like, you know, I don't know exactly what's going on in that class, but whatever this student was doing, they did really, really well.
A
Well, right.
B
And so, you know, it still doesn't solve the. Like, how do you reason about large versus small classes? You still need some information, right? Math 55 might have 10 or 20 students in it. And so someone has to know, oh, interesting. Like that class with 20 students in it, those 20 students all getting A's is meaningful, but at least towards this idea of like, very large courses, you know, with, with unclear grade distributions, maybe everyone gets an A. Or courses that are sort of known for being populated with a very large number of students who in one fashion or another are, you know, being elevated either either by the instructors or by the difficulty of the material or whatever. Like, seeing the grade distribution helps a lot. And again, remember, it's part of what has been hoped for, right? Like sort of what they've been trying to get, a broader grade distribution as Harvard. That's part of what the grade cap is sort of mechanically trying to enforce. But the difference is sharing the grade distribution, because it is directly diagnostic, actually reverses some of those effects we were concerned about before. Right. So if someone is taking a course with 700 students, all of whom get A's, that A, you know, just for the sake of having an A on the transcript, that A no longer looks as good. Right. And people have said, oh well, but like, you know, employers don't look at transcripts like, okay, sure, that's fine, but they could ask for this and, or we could provide this information as a component of GPA and expect people, you know, sort of in equilibrium to, you know, to choose to present. Oh, I performed extremely well relative to par in all the classes I was in. Transcript available. Right. So people can, once the information exists, people can choose to distribute it. And that means the, you know, firms and grad programs and, and professional schools can start to use it and even start to require it.
A
By giving the distribution, you give more context for the grade. But you're still not really solving for the difficulty of the course.
B
Correct.
A
You can infer some things about the difficulty of the course from some distributions. But. Okay, keep going.
B
Yeah, exactly. So it's not a perfect solution. Absolutely. And that's. And as I said, right, like you know, a 20 person, you know, maybe there's like a. But again, we'll stick with introductory Spanish and Math 55. Maybe they both have 12 students in them in one year. Like a 12 person Spanish section and a 12 person Math 55 section. And you're relying on the person reading the transcript to understand that even if in both of these classes all the students get A's, those things both have different meaning, Right. You have to have some semantic information of, okay, introductory Spanish labeled Spanish 1A. Like introductory Spanish is different from Honors Advanced Multivariable Calculus. I think it's called like, you know, sort of like, you know, sort of so many modifiers like math course. But I think that's not an unreasonable assumption, right? That people can, you know, be diagnostic using semantic information plus something about the grade distribution in the class. And that would get us a lot further. And in particular it helps with some of this like student filtering across classes issue. Because if your a in a 700 person class is going to be down weighted by your prospective employer, you don't have an incentive to take that class for the easy A anymore. Right? Like now, now we actually do get the selection effect we want. Right? You know, if the transcript reveals something about course difficulty and in particular it reveals when you're taking a very large, very high grade Sort of skew course. Now maybe people will go and seek out courses where, you know, there's sort of more heterogeneous performance. But at least, you know, sort of, you know, they're, they're not, you know, signal neutral in the way a large, a large easy A is.
A
So we, we, we are, by doing that in, in some way we are adding to the amount of signal because no longer are we only trying to get a single yes signal to do the work of two separate indicators, you know, or to, to do exactly indicate two separate things.
B
And, and, and there's, and there. It's not a perfect signal extraction here because again, the grade distribution isn't an exact measure of course difficulty. But we're giving you sort of two, giving the reader sort of two dimensions of information that speak to the student performance and the course difficulty somewhat separately. Exactly.
A
So if you wanted to maximize further, what would you do? And I know the answer because I've read the paper, but I'm interested in that.
B
Right. So this is, so this is where Josh Joshua Ganz and I say that you should use what we call Eigen grading. Oh, by the way, one other important thing towards your K12 point, you said like if only we could get all of the algebra courses to be consistent with each other. One other thing we show in the paper is that indeed in situations where they're sort of all the students take all the classes like you have in a lot of K through 12 programs there, you get more diagnostic, you can get more diagnostic information because you have a way of sort of like comparing students against each other, sort of like relative to each other other. Right? You see, you sort of see the relative performance of students in all classes. But the real problem, especially at the university level is when you start having sorting of students according to the difficulty of the class. Right. It's not just like, you know, it's not that students randomly select into advanced graduate material. It's the strongest students who've taken a lot of the like really hard undergrad courses then go on to take the really hard graduate course courses, you know, similarly like, you know, it's often not the people who are taking those graduate courses who are also going gem mining or whatever, right? It's like gems, remember from earlier as this, you know, sort of the, the easy courses. And so once you have a course that's really hard in whatever sense, it's often full of really strong students. And so you would most want to give those students higher grades and you'd want to account for the fact that that room was full of really strong students in figuring out how to balance grades across the university. And this is what we propose that an ideal system would be able to do. And we have a statistic, we describe a statistical method for doing it in the paper, what we call Eigen grading. As you previewed earlier. It relies on eigenvalues. But, but an intuitive way to think about is a little bit like the chess system. Elo. Right. So chess for, for those listening who haven't, who haven't come across this before, Chess has a ranking system. And the relative ranking of competitors is affected by their history of competitive play and sort of attends to in each match how well you expected them to do. Right? So when, when a player with a really high rating goes up against a player with a really low rating, if the high rating player wins, that's not surprising, right? That's sort of what we expect. And so their ratings don't move very much. Okay. The high ratings players rating goes up a notch. The low players rating goes down a notch because they won more, they lost more. But it's not a very big change because that's exactly what the ratings would have implied. They're high rated versus low rated. So we expect the high rated player to win. Right? But by contrast when you have the Goliath beat David.
A
Right?
B
Right, exactly. Yeah, yeah, yeah. Goliath's beat David most of the time. And so we shouldn't be surprised and so we shouldn't update our impression of Golia Goliath.
A
Right?
B
But when the David beats the Goliath, Right. So when the low ranked player beats the high ranked player, that's a big surprise, right? That doesn't necessarily mean the high ranked player is terrible. The entire history of their play was wrong, right? Maybe they just had a bad day, but it's still a surprise from the perspective of the rankings. And so the rankings move more. It's like a rubber band effect or like a slingshot, right? The, the high rank players ranking falls by more than it would have gone up if they had won. And the low ranked players rating goes up by more than it would have gone down. And so in that way the matches in effect account for the difficulty of the match in determining the impact on the, on the rankings. And so we imagine now, now chess is a competitive tournament sport. We're not, we're not saying that undergraduate college should be a competitive tournament sport. Just to be clear. The. But a similar sort of principle makes sense in grading as well, which is that, you know, we don't really know if we could, if only we could agree on a way of measuring course difficulty and like just put that on the transcript, that would completely solve the problem. Right. Once you can write down exact course difficulty, you're done. The only problem is it's not even clear how you would measure that. Right. Like, much less getting all the faculty members to agree on it.
A
Right.
B
Like, you know, anybody who's, who's been in a, you know, in a faculty meeting knows that it's very difficult to agree on much easier things than who teaches the easy classes and who teaches the hard classes class. Right.
A
Quite a political contest. But if you use a bunch of students and their relative grades, all of a sudden you have dimensionless difference that you can reference to, right?
B
Bingo. Exactly right. And so like what is a hard class from the perspective of grading, you know, and what is a strong student? Well, in the same way that like a chess, like a high ranking chess player is a player who, if you put them up against almost anybody, high ranking or low ranking, you expect them to perform well. A high ranking student is a student who, no matter what classroom they're in, they do well. Right. And so the idea of the metric is a strong student is a student who takes many classes with other strong students. And what is a strong student? It's a student who, you know, has high performance no matter what classroom. So whether they're taking math 55 or econ 5067 or whatever, like they're, you know, they always do very well. And so then if you have a room where all the students are students who, no matter where they are, they always do very well, you treat their performance there, it's like, it's not surprising if they all do well there. Right. Like everyone getting A's. If the students who get A's, no matter what class they take, all get A's in a class together, we shouldn't be surprised. That's just them performing like they perform. Whereas if you take a student who underperforms in most classes and put them in a room with a bunch of students who get A's everywhere, and that student who normally underperforms does really well, that's a big update. Like, whoa, they're really good at this thing. That grade matters a lot. And so you can imagine re ranking GPAs. I've said to sort of describe this as like a static process, but you can use the information of who takes which courses with whom and what their grades are to establish sort of like an overall ranking that implicitly differences out information about course difficulty.
A
And so for the admittedly small portion of our audience that understands sort of IRT equating on test scores, you have a bunch of kids who are taking some of the same questions, they have different difficulties, and then you can identify where the students are by comparing questions with different difficulties. This is the same thing to a degree, but with courses. The only thing is that there's also some political dimensions about, well, which courses should I take and how difficult should I make these courses that aren't in the situation of an assessment where you would use irt. So if you could do this with the eigenvalues, what you would essentially be doing on the transcript, which is kind of important because it's not like waving a magic wand, it's saying, well, we could provide more information on a transcript and if we could understand how difficult the course was, we could both give a grade and a difficulty. And theoretically we could use both of those pieces of information for a grade point average that's adjusted. That takes away the differential difficulty of courses. Do I have this right?
B
Yes, absolutely. And with the additional wrinkle that when you can't measure the difficulty directly, using inference across students, performances across classes is sort of a, like a, like a best substitute. Right? It's the, it's a way of inferring the difficulty statistically, even if you can't measure it directly.
A
So what does this do to those, those sort of political pushes or polls, or let's just call them the incentives on the part of professors. I mean, to some degree, no longer do they get a lot of benefit from having a more lenient course. Because now we have a number of how that indicates how lenient their course is, which might be something they don't want displayed, but also seems like the students realize, well, hold on a second, my A in that class is not going to be worth as much if my A and this sort of alternative to the GPA factors out, or basically I get a discount based on the, the difficulty of the course.
B
Yep, you've completely nailed it. So the, the idea is that anything you can do to make the information value of the grade more sensitive to the, to the actual difficulty of the course, like the downstream information. Right. The impact of this on, on how employers and, and, you know, professional schools and whatever will view the student's transcript, the more you can encourage students to start sort on the basis of, you know, sort of what they actually want to take rather than whether they can get an A because, you know, sort of like Just even, even in that simple example with the 700 students all getting an A, right? Like that A on a transcript becomes much less meaningful. Right? Once you, once you can't hide the fact that everyone in the class, it was a huge class and everyone got an A. And with Eigen grading it gets even better. Right? Like now when you take a course that is sort of easy in the sense that it's not very diagnostic of performance given all the courses everyone's taking and what they're, you know, sort of, and what their sort of equilibrium overlaps are, there's no benefit from the student's perspective to taking the easy course just for the sake that it's easy. If they want to take that class, great. But like in terms of the transcript, they don't gain anything from it. And so unlike the grade cap, which as we talked about has this, you know, sort of very perverse unraveling incentive to push people away from the harder courses under something like Eigen grading, there's actually the reverse. It's like if you want to take a harder course, like the, the GPA you get reflects the fact that you enrolled in this course that is much harder than average.
A
And I really, really like this part of the paper and I like it because we talk a lot, particularly in higher ed about sheepskin effects and well, you know, we just want to get this thing that will make us look better in the market downstream and that seems like a sort of a non optimal way to set things up.
B
Up.
A
Let's set things up as a rigged game so we can maximize that. Seems bad. Seems like instead we should say, you know, you should be engaged to learn things. And to the degree that we can set up both transcripts and grades so that they incentivize people to think about what they want to learn and to work at learning that, that that's a much more optimal system. And so that's what I think is so interesting about this. Admittedly complex, it's going to be a little bit difficult to put this into operation. But let me ask you about that next because I really do think that that part of the incentives package where students all of a sudden would be motivated not to prioritize the easy A, but prioritize the pedagogical value of what they're taking for a near term step to do this, what would you need, you'd need a column on a transcript and an alternative way to sort of say, well, this is a different kind of measure like a GPA and you'd need Transcript data and grades right now, you wouldn't need to share it, but you'd need to be able to crunch the numbers.
B
Yes, that's right.
A
So how hard would this be? And I know that you don't work for a college administration, but I mean, how much of a lift is this? For instance, I asked this because so often we hear people say, well, I wrote an academic paper and it's a crazy idea. It's so far from reality. In some sense. This doesn't have to be that hard. Or am I reading it wrong?
B
Yeah, I mean, I certainly think this is implementable. A lot of the challenges actually come from things about stationarity, right? Like, how does it, how does, you know, how do you deal with the fact that students are in university for multiple years? And so the, you know, sort of the imputed difficulty of a class might change depending on who took it in what year or who taught it or something. So there's a lot of questions at the implementation level about sort of stationarity of the distribution and other statistical properties. But from, but, but I think those are surmountable. And I think it's very much implementable. I mean, like, from a, you know, from an algorithms perspective or whatever. Like, we, we know how to compute, you know, spectral, like, what this whole class of what are called spectral measures. The, in fact, we, you know, you know, Josh, Joshua Ganz and I have done simulations and stuff. Like, it's eminently doable. And by the way, it's similar to the things we use like for PageRank on the Internet. Like, there's a lot of other types of measurements that have these flavors. I think there is a challenge with, you know, sort of with explaining anything that's got a, you know, complicated statistical under sort of underpinning, you know, and making, making it legible. Right. Like, one thing, as I said, you know, a feature of the grade cap is it's extremely legible. It is easy to explain. You can say it in a sentence, 20% plus four A's. Like, you know, we're doing a statistical reweighting to try and like, you know, better, better improve the information in grades, the information content of grades. That requires sentences and probably draws at least a couple question questions. But I think it's eminently doable where we're, you know, as a side note, anybody who listens to this and thinks, wow, we want to try that or we want to learn more, please feel free to reach out. And the other thing that I'll say is that the general principle is Eminently doable, right? If the, if you take the principle as the way to reduce grade inflation is to change students incentives to seek the easy A. And the way to seek, you know, change students incentives to seek the easy A is to provide information that makes the easy A less valuable for students. As I said, providing course grade distributions, median scores. One of my, one of my friends, David Vendler, said that when he was in college they, everyone got a couple of sentences describing like, sort of like, like very, very short but like informative, like information about their performance attached to the grade and like that is another example, right? If like, you know, the instructor can't find anything positive to say about the student, maybe that's indicative, you know, sort of anything that changes the information content in a way that reduces students incentives to sort on course easiness and instead, as you said, like sort of drive students to sort on the basis of what they actually want to study.
A
And so, I mean, some of this is sorting. But before we end, I just want to get back to it that on the one hand, I understand the advantage of the Harvard grade capping approach. It is very sensible, it's very intelligible, it makes sense, it certainly solves the public relations problem. But what it doesn't seem to do is push students to work harder. It doesn't seem to give them, you know, getting incentives right is a big part of the whole ball game. And so, yeah, to the degree that you could read a transcript, a transcript system and a grade system that would push on those incentives in the right direction, I think that more and more thoughtful information along these lines where you try and separate course difficulty from grades and really try and get that one indicator to do both things at once rather than to try and be an indicator of two things makes a lot of sense. Scott, let me ask you one question on the way out, and this is just bigger picture perspective. We had a Yale report that was kind of a mea culpa on grade inflation for the Ivies. We have the Harvard action here. But when I talk about these incentives, a lot of the incentives are still in place, not just at the Ivies across the board. Same in K12. What do you, what does your gut tell you about this sort of great inflation reckoning? Is this a turning point or should we not be too hopeful?
B
It's a great question. I mean, and I don't have a crystal ball, and this won't be a surprise to you and it won't be a surprise to your listeners, but markets have this tremendous Ability to respond to incentives. Right. And so, and, and when you don't get the incentives, right, the system isn't stable, right. Like it sort of, it, it pushes, it's like a, you know, squeezing a balloon or something like this, right? You're like, you know, you're pushing, you know, you're sort of pushing in a direction that the market is trying to take you. And one thing actually that we didn't get into, right, is like, well, on the one hand there's this view that the downstream grades need to be more diagnostic for people's, you know, professional aspirations or whatever. There's also the view that, you know, jobs are getting more competitive, it's getting harder and harder to get them. And so towards students, professional aspirations, especially in a world where grade inflation is happening elsewhere as well, perhaps students are going to face more and more pressure to seek the higher grades. And so if we create a system that puts them in this sort of, you know, rat race competition with each other to sort of like make sure that you're ahead of your classmates, that actually has the tremendous potential to backfire because the downstream incentive hasn't changed, right. The need to get a high grade, you know, for, you know, for distinguishing yourself in the workplace or in, you know, the professional marketplace, like that hasn't changed. And so there will be tremendous market pressure, you know, sort of inside the grade system. Right. This is not a closed environment, right. This responds to these outside, you know, sort of market incentives and then it'll, you know, create, I think, a lot of pressure to unravel and to seek easy courses and lead to something that's likely to be pretty unstable. I mean, I hope I'm wrong, but like, I really think you, you know, and the lesson for market design over and over again as a field is that you can't really solve a market failure unless you redesign the system in a way that aligns the incentives. And so I think at some fundamental level we don't get, we won't really have solved the problem of grade inflation until we find an incentive aligned solution.
A
Thanks for listening to the report card with Nat Malkus. And special thanks to our guest Scott Duke Commoners. Also thanks to our producer Ellie Lucas. He makes this podcast possible. That's all for this episode. I nap mountains.
Episode: What Harvard's New Grade Inflation Policy Gets Wrong (with Scott Duke Kominers)
Date: June 3, 2026
Host: Nat Malkus
Guest: Scott Duke Kominers, Professor at Harvard Business School
This episode explores Harvard's recent response to grade inflation—capping A grades at 20% per course—and why economist and professor Scott Duke Kominers believes this policy is misguided. Drawing on his research and practical experiences, Kominers discusses the problems with grade inflation, the unintended consequences of grade caps, and presents alternative solutions that could more effectively align incentives for students and faculty. The conversation delves deep into the information content of grades, student sorting, faculty grading practices, and innovative proposals for maximizing transparency and fairness in academic grading.
"It is absolutely necessary that we stop grade inflation. Grades are out of control…Employers have trouble interpreting Harvard grades…students feel it is necessary to get a perfect GPA…" — Kominers (playing devil's advocate) ([18:48])
"Your ability to get the top grade depends on how you do relative to the classmates sitting next to you…now you’re competing against your classmates." — Kominers ([29:00])
"If you see that someone got an A…in a class with 700 students and 700 of them got As, you’re starting to form an inference about the difficulty of [that course]." — Kominers ([54:22])
What Is Eigen Grading?
Application and Benefits
Practicality
“If we could provide more information on a transcript…We could both give a grade and a difficulty [score]…that takes away the differential difficulty of courses." — Nat Malkus ([66:30])
"You can’t really solve a market failure unless you redesign the system in a way that aligns the incentives. I think…we won’t really have solved the problem of grade inflation until we find an incentive-aligned solution." — Kominers ([77:50])
On Holding High Standards
"Holding students to a very high standard is incredibly important for establishing the trust relationship...it's not because you want to see them fail, but because you know they can succeed." ([04:20])
On Grade Caps
"When you impose a grade cap, you actually are regressive in terms of your ability to reveal this information." ([24:04])
On Student Sorting and Perverse Incentives
"There's a lot of risk now. When you take an advanced class, you are putting your grade at risk if you stretch yourself." ([33:30])
On Solutions
"Providing course grade distributions...changes students' incentives to seek the easy A… now, now we actually do get the selection effect we want." ([58:00])
Scott Duke Kominers offers a rigorous critique of Harvard’s new grade cap, highlighting its regressive effects and misalignment of incentives for both students and faculty. He proposes actionable, transparent alternatives—beginning with grade distribution disclosure and culminating in Eigen grading—designed to restore the informativeness and fairness of grades, and to encourage genuine learning and academic challenge. Ultimately, Kominers argues that only systems aligning market incentives with educational values can stabilize and improve higher education grading practices.