May 28, 2020
Derivations of the Fisher Information
Some theory, some examples, and some insight

By Andrew Rothman
2 min read
Background and Motivation
The Fisher Information is an important quantity in Mathematical Statistics, playing a prominent role in the asymptotic theory of Maximum-Likelihood Estimation (MLE) and specification of the Cramér–Rao lower bound.
Let's look at the definition of the Fisher Information:
The descriptions above seem fair enough. The Fisher Information is the variance of the score. Simple, easy, great!
Textbooks often state (sometimes without proof) that under regularity conditions the following three quantities are all equal to the Fisher Information:
When I first encountered this material as an undergraduate, it wasn't clear to me how these three quantities all held equality? A quick search on Medium revealed a lock of coverage on this topic. I think the proofs below are worth knowing. Some of the techniques used in these proofs are useful elsewhere in Probability Theory and Mathematical Statistics. So let's dive in.
1) Fisher Information = Second Moment of the Score Function
2) Fisher Information = negative Expected Value of the gradient of the Score Function
Example: Fisher Information of a Bernoulli random variable, and relationship to the Variance
Using what we've learned above, let's conduct a quick exercise.
Final Thoughts
I hope the above is insightful. As I've mentioned in some of my previous pieces, it's my opinion not enough folks take the time to go through these types of exercises. For me, this type of theory-based insight leaves me more comfortable using methods in practice. A personal goal of mine is to encourage others in the field to take a similar approach. I'm planning on writing based pieces in the future, so feel free to connect with me on LinkedIn, and follow me here on Medium for updates!