Learning one thing makes a network forget another.
A neural network learns by adjusting shared weights. When it learns from a stream of experience, as a reinforcement-learning agent does, an update that helps today can quietly damage what it learned yesterday. This is called interference.
Interference is worst when every input lights up most of the neurons (a dense representation). If each input used only a few (a sparse one), learning about one input would mostly leave the others alone.
Fuzzy tiling activation (FTA), from Pan, White and White, makes a layer sparse by design. The catch, which this paper tackles, is that it only works if one number, the tiling bound, is set right.
How FTA works. Move the sliders.
ReLU turns a number into another number. FTA turns a number z into a vector of k tiles. It splits the range −u to u into k tiles and lights up the one z falls in. Neighbouring tiles light up partly, fading over a distance η (the "fuzzy" part), so the network can keep learning through them.
Set the bound wrong and the tiles go to waste.
In a real network, FTA's inputs come from the layer before it, and nobody knows their range in advance. Here a layer produces 64 values with a spread you choose. Pick a bound for plain FTA and count how many of its 20 tiles get used. The right-hand card always uses tanh with u = 1.
What the paper did.
Reproduced the original result
A deep Q-network (DQN) learns to land a rocket in LunarLander. With FTA in its last hidden layer and u = 20, it matched the pattern the original FTA paper reported.
Showed FTA is fragile
Sweeping u from 0.1 to 100, FTA did best near u = 1 and badly when u was far off. The best u also differed between LunarLander and CartPole, so it has to be searched for each task.
Found a fix: tanh
Of tanh, Batch Norm and Range Norm in front of FTA, only tanh worked at every bound tested in LunarLander. It also fixed CartPole's bad bounds, and its results were the most consistent across random seeds.
The evidence, figure by figure.
The paper's six figures. LunarLander runs last 500,000 steps and are averaged over 10 runs; shaded bands are 95% confidence intervals.
Re-run from scratch
Every hyperparameter is a search someone has to pay for. tanh squeezes FTA's inputs into −1 to 1, so u = 1 always fits, and a sparse activation works out of the box.
Limits: two environments, one algorithm (DQN) and 10 runs per agent. Whether tanh-normalised FTA holds up across more tasks is open.