Nicole Junkermann

Ideas and technology

Nicole Junkermann on Reinforcement Learning, Explained Calmly

An idea that fits in one sentence: try something, take a signal, adjust. Nicole Junkermann tells it slowly, with a kitchen analogy and no promises.

Nicole Junkermann smiling in a bright portrait with pale curtains behind her
Learning by repetition has more of the craft about it than the miracle, writes Nicole Junkermann.

Reinforcement learning sounds technical and is a good deal less so than it seems. The underlying idea fits in a sentence: a system tries something, receives a signal telling it whether that went better or worse, and uses the signal to adjust what it does next time. Repeated a great many times, the cycle produces sensible behaviour without anyone having written every step down in advance.

Learning the way you learn to cook

The easiest analogy is cooking. Nobody learns to cook by reading alone. You try, you serve, somebody pulls a face, and the face is worth more than half a recipe. In time the hand knows how much salt before the head has reasoned it out, and it knows because it has been collecting small corrections.

A system learning by reinforcement does something comparable, without taste and without hunger. What it receives is not a flavour but a number, and what it adjusts is not a hand but a preference for certain actions. The mechanism differs; the shape of the curve is much the same.

It helps to hold on to the smallest parts, and there are only a few of them:

  • An environment where things happen and where every action has a consequence.
  • A set of possible actions, which may be four of them or a very great many.
  • A signal saying whether the result came out better or worse than before.
  • Many attempts, because a single try teaches absolutely nothing.

Why patience is part of the method

Here is the part that most resembles ordinary life. This kind of learning is slow by design. It has to get things wrong many times over to tell luck apart from judgement, and it needs somebody to have chosen well what gets rewarded, because a badly chosen reward teaches precisely what nobody wanted taught.

There is also a delicate balance between repeating what already works and trying something new. Staying with the familiar gives acceptable results straight away; exploring costs at first and sometimes pays later. Anyone who has picked a restaurant in an unfamiliar city will recognise the dilemma without needing it explained.

That is perhaps the detail most often lost in the telling. There is no shortcut. Improvement arrives by accumulation, and accumulation asks for time and a certain tolerance of error.

How far the metaphor goes

Not very far, and it is worth saying so. The method understands nothing of what it does, holds no intentions and is not good for everything. It is one concrete way of tuning a behaviour through repetition and signals, and within its own ground that is already useful enough.

Put that way, with the noise removed, the subject loses mystery and gains interest. Nicole Junkermann prefers that register for anything technical, the one argued for in Escribir when something difficult has to be explained without being flattened. The previous entry, on artificial intelligence in India, looks at the same technology from outside, and the rest of the notebook continues in the English notes.