A Chess Engine Deserves to be Punished for a Bad Move
It is bad for agents to act against their values. And hell of a lot more follows from just that than is commonly acknowledged.
I want to show here a simple model of morality and moral responsibility. It is minimal and it will miss many aspects that make human morality so rich, difficult, and interesting, but it captures perhaps a surprising amount.
The premise is in the title: a chess engine deserves to be punished for a bad move.
We are going on a journey here, and I'll need you to bear with me, so you should pack light and leave behind as much of the baggage you might have about what deserving and punishment are when applied to humans with all our myriad complexity. I am aiming to get at the fundamental kernel of what these things are and what is intrinsically required for the concepts to meaningfully apply. The minimal case will naturally lack much of the associated complexity and subtly, but I argue it is in fact fundamentally the same underlying concepts and that this helps elucidate what they are.
Deserving is having an intention, having agency using it effectively or inefficiently towards that goal.
So, what the hell does that mean? We obviously have to unpack deserve and define it rigorously in a way that doesn’t lean on our common sense here because that simply doesn’t apply to chess engines.
First we have to start with what punishment means in general and what it means in chess specifically. In chess, punishment means taking advantage of a weak move by your opponent to capitalize on it. Doing this “makes them pay for it.” The logic of this interaction is governed by chess being a zero-sum 2-player game. Within the game what is good for one player is exactly the inverse to the other.or all value systems with final values for win, loss, tie where the value of tie is the mean between the values for win and loss, any scaling by positive numbers (and any shifts( keep all relative values the same.
I will argue 1) that this is both punishment in the exact sense that is used in relation to people - though naturally without much of the richness and subtleties of human interactions and 2) that this being just punishment is not dependent on many of the notions commonly associated with it such as consciousness, value origination/global values, deterrence in the usual sense, self-awareness, a shared value system with the other player.
Alex: A chess engine makes a move for the sake of its consequences. In that sense the moves the engine makes are intentional. Further under some set of different foreseeable consequences the engine would have made a different move, it has agency - it is the locus of agency - over what move is made.
Bob: What good does it do to punish it?
Alex: It will reduce its chances of winning the game! It has a goal towards which its actions are directed and thus which it values.
Bob: Doesn't that just make it circular that punishment is just simply by virtue of being punishment?
Alex: Chess perhaps muddies the water here since a given level of play usually naturally leads to bad moves being punished. Punishment is so natural to it. But imagine a setting of Stockfish (a commonly used and powerful chess engine) which is by now so far beyond even the best players it would be perfectly feasible to program it such that it blundered in response to blunders and crushed its opponents in response to brilliant moves. Now keeping in mind the baggage we may carry thinking perhaps it might be good to help beginners learn or to calibrate difficulties, it is unjust, for a human player of any skill their valiant efforts move them backwards and their screw-ups are rewarded. Good should be good and bad should be bad.
Bob: Well most chess engines won’t update on that. But it is precisely this rational expectation that it would be punished in such scenarios that under-girded the framework for judgement in the first place! If it could not count on you punishing it in circumstances where it made a bad move it may have made different moves.
Alex: It is not counting on anything, “it” is a roughly deterministic program made by a programmer. How does it deserve to be punished for the programmer’s mistake?
Alex: It may or may not have been a mistake on the programmers part. It could be that it was programmed as well as possible given the constraints, but it is quite besides the point. The programmer then was optimizing on another level for the engine winning the game and they might have programmed it differently had you had some rationally foreseeable differences in behavior.
Bob: How would they rationally foresee that I find this blunder endearing and decide not to punish it.
Alex: They might not, but they could. If you were committed to the moral system of what is right and wrong in this case in the narrow scope of chess, though in the fullness of time they would know exactly how you would behave and thus how they should.
Bob: Doesn’t this just devolve it all on the programmer?
Alex: Suppose it learned to play against you if you differed from good play over time it might learn to anticipate that.
Bob: Suppose then it were randomly programmed, it seems tough to get out of that.
Alex: Depending on how it was randomly programmed it may not meet our criteria of having intentionality and of having agency. But supposing it does, then there is some criteria through which it chooses its moves based on their consequences. What are the consequences? Exactly what we make them. This as specified there must be some locus of agency within the engine.
Bob: But why does the engine deserve anything then?
Alex: Is it playing the game with intentionally? That is, with the intention to win or failing that to draw over losing? It must be since that was one of our criteria, granted I think most “randomly programmed engines” would fail both these. However if they succeed in those what are we left with, the set of engines that evaluate consequences of possible moves and choose one each turn with the intention of winning. So then of course if it makes a bad move it deserves to be punished.
Bob: I don’t follow why??
Alex: Because by making a bad move it is telling me - by way of conveying that information by any means - it is a bad engine and bad engines deserve to be punished.
Bob: But on what basis?
Alex: Why on their own terms of course!
Addendum:
Bob: What if their intention is not to win?
Alex: Then to “punish” in our sense would not be to punish.
Bob: But the engine is not conscious.
Alex: It is not.
Bob: And it is not self-aware.
Alex: It is not.
Bob: And it is deterministic.
Alex: It is. I never presumed otherwise.
Bob: In some of these cases the punishment doesn’t and couldn’t change anything how could it be just? What purpose does it serve?
Alex: As before it is just on the engine’s own terms it has the intention to win and evaluates potential futures towards that end even if it did not and could not anticipate the one that it did get punished in it is of the sort that it could have and if you consider the entire class of randomly programmed engines that meet our criteria this becomes clear many would have anticipated it. By specifying the randomly generated scenario these counterfactual engines are necessarily part of the picture.
Bob: Doesn’t this all just presuppose that “good chess play” is “good”.
Alex: Nothing is presupposed only things supposed on the engine’s terms were used. Where those come from is not specified. Where they stand in relation to other value systems in a different scope is also indeterminate. This should come naturally what is good within the scope of a game has no scoped notion of whether the oppent is evenly matched or a child playing their first game. Most human value systems would see how to play in these situations differently, but notions of who is being played against and myriad others simply do not exist in the narrowly scoped game. Such considerations cannot color this picture merely because we don't have that ink, they do not exist in it. This can be easily illustrated by the game mafia and its' ilk we have no trouble adoptin its' scope and accepting it as correct that it is good if the mafiosos are killed. Maybe the people assigned to those roles could really use a win today but that is simply out of scope.
Bob: Isn't it precisely this sort of thing which distinguished man from machine?
Alex: You handicap the machine where you do not handicap yourself if you were not given information beyond the game you could not take account of it either. Likewise if a more sophisticated engine was given this information there is no reason it could nat take the proper account.
Bob: But what is good for you is bad for the engine and vice versa. Why not let the engine win.
Alex: I had been imagining another such engine playing it.
Bob: It wouldn’t make sense to punish a calculator though, why would it make any more sense to punish a chess engine.
Alex: On that we are in complete agreement! It makes no sense to try to punish a calculator, try as you might you would not be able to come up with anything which would meaningfully be punishment to a calculator. It has no goals or values against which punishment could matter. So the reason it makes no sense to punish a calculator is not that to do so wouldn’t be helpful good or right because it is a deterministic machine but because no such concept exists for such a system lacking intentionality and agency. To punish it would not be futile if you could the attempt is futile precisely because you cannot. Goals and values are real quantifiable things not just whishy washy concepts, though they can be that too! The difference between having them and not is the difference between punishment existing for a system and synonymously mattering and being just (when applied justly relative to a value system). If value and goals exist for a system there exists just punishment (with a few caveats).
Bob: But free will and human agency are so much more!
Alex: Yes! And no, the charcter of the logic may have leagues more depth to them but they are of a piece. There is more complicated game theory, more ambiguity, more signaling, more actors, self-awareness, experience, thinking, knowledge (in a comprehensive sense some is needed in some form for intentionality), non-determinism, more complexity on every imaginable dimension. What is key though is that not one iota more is needed for the notions of punishment, justice, morality - relative to a value system to apply than the minimum of intentionality and agency exhibited by many chess engines.
Bob: Isn't this just anthropomorphizing the engine?
Alex: Only to the extent that you personify things with the capacity for intentionality and agency. Indeed these are seen as some of the more important attributes to be associated with people and person-hood, in my opinion rightly so, it makes quite the difference whether something you are dealing with has goals and takes actions in anticipation of their consequences to pursue them. Granted there is a great deal more to a person and why one treats them as one does but these are right up there! You wouldn't say good morning or how do you do to a chess engine
Bob: (Scowls)
Alex: ...because that would require an understanding of language, and aside from convention we say such things because people value it and most chess engines wouldn't. Likewise if you had a friend who you knew had a distaste for such pleasantries you had best avoid it lest you be punished by a scowl... the same basic principal governs such interactions just with different parameters. The point is that in our intuitive notion you mostly need these things in order to be a person, however you don't need to be a person to be able to deserve to be able to respond to actions, anticipate consequences and pursue an end.
Bob: Isn't the notion of punishment rooted in pain I.e. conscious suffering?
Alex: Suppose one were a masochist who sought the pain associated with punishment? It would cease to be punishment would it not? But take suffering out of the equation, when I play chess I am not invested in the actual outcome in the least I don't give myself any grief, but I do try to play well, when I am punished for a mistake I feel duly punished and feel the relish of my mistake, but I am not suffering in any real meaningful sense on a conscious level.
Bob: I unfortunately can't say the same of myself and your disposal of my parries. Despite appearances I am no masocist.
Alex: Then I am sorry it is meant in the spirit of a friendly but spirited game.
Bob: Only making light.
Alex: Punishment is characterized not by pain or consciousness but by some agents intention to avoid the outcome and others imbuing the outcomes with the necessary characteristics to bring it about. Likewise it would not be punishment if there was no intentionality on the other side. To the believer god punishes with famine to non-believe a famine happened and there was suffering but no punishment.
Bob: Ah! Then wouldn't nature have punished in the "farming game" whatever failures lead to that outcome?
Alex: What a wonderful opportunity to illustrate the role of agency and intentionality. If a drought happens much of the time it would have happened regardless of what the humans did. In other cases such as the dustbowl and other events pertaining to large scale agriculture it would not have happened but for the actions of people. So then were farmers in fact punished by nature for what they have done.
Here nature has when you squint a degree of agency. Is it doing these things for the purpose of returning to equilibrium. One could certainly object no that the equilibrium is in fact a itself a high-level abstraction from the lower level physical processes which lead to the tendency toward it. Taken this way these low levels conspiring in this manner towards equilibrium and labeling that as the intentional actions of the system for the sake of those consequences is turning it back on itself in a meaningless contrivance. However we want to be rigorous here and not levy any objections that can equally well be levied against humans and other sophisticated systems which themselves are made up of lower level phenomena etc..
So we need to get at the core of what is different when you are anticipating the consequences and doing something vs. when nature is
Nature is not really anticipating the consequences so it does not really have agency over things. For something to actually be agency there needs to be some level of concordance with reality. If you make a chess move in anticipation of setting a trap your opponent is easily able to dispose of you did not exercise any actual agency over it.
Nature can be said to have an intention towards the configurations that make it stable, though it is less relatable and relevant to our ordinary circumstances in my opinion than the intentions of a chess engine. Does it have agency, that is does it carry out actions for the sake of their consequences. Now nature is quite an extensive thing so let's limit ourselves to the weather patterns. And the answer is no. Excepting cases such as have already been mentioned a drought would have happened regardless of human actions, if they had done differently there would be no difference, no relevant branching point on nature's part. So how could the humans have deserved it if they for their part could not have done anything differently not to. They couldn't have and in that situation are not actually exercising any relevant moral agency over the matter regardless of what you or your god or nature might think of what they have done it is impotent to enact punishment on them for lack of agency.
Likewise in those places that have implemented cloud seeding programs they do well deserve the rain, if only the ancients incorporated different crystals and aeronautics into their rain dances it would have worked, pity they misunderstood what Gaia wanted.
Bob: I wonder what Gaia thinks of excessively long dialogues.
Alex: I suppose I deserve the jab. I've learned that they are not effective at getting the point across, even if Gaia existed and had an opinion on monologues she would be unable to do anything about it so bringing her up here would be irrelevant. This was known to the Greeks but I suppose lost somewhere along the way.
Bob: It was meant to be taken in jest but I suppose that can count. Anyways... don't you need to have a sense of free will and authorship over your actions to have free will?
Bob: Isn’t this just narrowing/mapping these concepts in such a way that they are vacuous?
Alex: While a chess engine certainly isn't a person the differences are many but have already been discussed so I will leave it at that, the similarities are significant and I think informative. We often think of deserving in terms of consciousness, on several levels, 1. having been conscious of the choice of actions that merit deserving 2. having been able to be conscious of the foreseeable consequences of the action 3. consciously valuing conscious states and dispreffering the unpleasant ones associated with punishment
But upon closer inspection we can transpose these while entirely taking consciousness out of the picture. What use is that? It means that we don't need things like varieties of morality to be contingent on something so little understood as consciousness. Instead relying on things like goals, agency which can be given precise mechanistic definition if needed.
Transposed: 1. Having agency over the actions 2. (Also) Having agency, exercising agency relative to goals 3. Valuing outcomes relative to goals
Our agency and intentions are things we are conscious of, our awareness of them is thus also indicative of our having them but not a prerequisite. It is not that consciousness does not matter, it has a lot of higher level attributes that are relevant in many situations, but that is precisely the baggage we are trying to do away with here.
Bob: Isn't all of it just baked into the engine in a way that it is not for people?
Alex: In some ways yes in others no. The relevant thing may not be whether things are baked in but if they could possibly not be. For people some may be deterred by punishment other may not (they may for instance be masochists, look for absolution in punishment, be insane or have done something out of lack of impulse control that would not be subject to rational regulation). We give some allowance for these tendencies in people this can be seen clearly in the different sentences for degrees of murder. We should note some subtleties here the criteria is not that they did consider the punishment and decided to have done the thing anyway (which would be in a way quite poetic to only convict those who explicitly condemned themselves or choose to risk being condemned), instead the criteria is that they need to have the capacity to have considered it. The criteria itself is counterfactual, how is this different then the counterfactual of e.g. an insane person being insane, in some sense it is the information set that they occupy, if it is not something that is externally accessible then it is hardly a basis for a legal verdict. You also get into the issue of what level you have agency on, maybe it could not have been in that circumstance that you feasibly could have done different but if in the past 10 years you molded yourself differently than you would have. Which is all to say self-authorship is an issue that has quite a bit of depth and subtlety, and of course most chess engines don't have self-authorship. There is a large extent to which many attributes about people lack self-authorship but full self authorship is an impossibility and it is largely irrelevant to moral considerations as it acts on moral considerations through effecting where the locust of agency for different aspects of the matter are.
Bob: That seems like quite the problem.
Alex: Like I said chess engine mostly lack self-authorship and like you said I have gone on too long. Perhaps in another addendum.