Introduction#
Johnson noise was first discovered by John Bertrand Johnson at Bell Labs in 1927 when he was investigating the causes of background noise in vacuum tube amplifiers. He set up several experiments to try to isolate the source of the noise, and discovered that the noise was due to random voltage fluctuations on conductors. He noticed that it was independent of frequency (today known as white noise), and it was directly correlated with temperature. In short order, Harry Nyquist provided a theoretical solution to the problem in 1928 in his paper titled “Thermal Agitation of Electric Charges in Conductors”. It is the result of this paper that we are going to set up from first principles to show how absolutely insane of a feat this accomplishment really was.
First, the conductors mentioned really mean all conductors, including resistors. Because these fluctuations are electromagnetic in nature, Nyquist had the brilliant idea to set up an experiment using two resistors of equal value $R$, and he connected them with a transmission line with characteristic impedance $Z_o = R$. The setup is shown below:
The premise is that each resistor is acting as a voltage source, and the electromagnetic waves propagate from one resistor down the t-line, at which point they terminate in the opposite resistor. We will now instantiate the mathematical construction of this setup.
1. Standing Waves on a Transmission Line#
Each traveling wave component (forward wave $V^+ and incoming wave $V^-$ is represented as:
$$ V^+(x,t) = V_0 e^{j(kx-\omega t)} \quad and \quad V^-(x,t) = V_0 e^{j(-kx-\omega t)} $$Combine for total wave expression as a function of time and position $f(x,t)$:
$$ V(x,t) = V_0 e^{j(kx-\omega t)} + V_0 e^{j(-kx-\omega t)} $$Using Euler’s formula $e^{j\theta} = \cos\theta + j\sin\theta$ and substituting, we arrive at:
$$ V_0 e^{j(kx-\omega t)} = V_0\left[\cos(kx-\omega t) + j\sin(kx-\omega t)\right] $$and
$$ V_0 e^{j(-kx-\omega t)} = V_0\left[\cos(-kx-\omega t) + j\sin(-kx-\omega t)\right] $$Summing the two wave expressions, we get:
$$ V(x,t) = V_0\left[\cos(kx-\omega t) + j\sin(kx-\omega t) + \cos(-kx-\omega t) + j\sin(-kx-\omega t)\right] $$Recall the angle-sum/difference identities:
$\cos(\alpha-\beta) = \cos\alpha\cos\beta + \sin\alpha\sin\beta$
$\sin(\alpha-\beta) = \sin\alpha\cos\beta - \cos\alpha\sin\beta$
$\cos(-\alpha-\beta) = \cos(-(\alpha+\beta)) = \cos(\alpha+\beta) = \cos\alpha\cos\beta - \sin\alpha\sin\beta$
i. (cosine is an even function from its power series: $\cos(-\alpha)=\cos(\alpha)$)
$\sin(-\alpha-\beta) = \sin(-(\alpha+\beta)) = -\sin(\alpha+\beta) = -\sin\alpha\cos\beta - \cos\alpha\sin\beta$
i. (sine is an odd function from its power series: $\sin(-\alpha)=-\sin(\alpha)$)
Expand and combine like terms:
$$ V(x,t) = V_0\left[2\cos(kx)\cos(\omega t) + \cancel{\sin(kx)\sin(\omega t)} + \cos(kx)\cos(\omega t) - \cancel{\sin(kx)\sin(\omega t)} + j\left[\cancel{\sin(kx)\cos(\omega t)} - \cos(kx)\sin(\omega t) - \cancel{\sin(kx)\cos(\omega t)} - \cos(kx)\sin(\omega t)\right]\right] $$Some terms above cancel, so we will rewrite the simplified expression:
$$ V(x,t) = V_0\left[2\cos(kx)\cos(\omega t) - j2\cos(kx)\sin(\omega t)\right] $$Thus
$$ V(x,t) = 2V_0\cos(kx)\left[\cos(\omega t) - j\sin(\omega t)\right] $$By Euler’s identity, $\cos(\omega t) - j\sin(\omega t) = e^{-j\omega t}$ (where $\theta = \omega t = 2\pi f t$), so:
$$ \boxed{V(x,t) = 2V_0\cos(kx)\,e^{-j\omega t}} \tag{1.1} $$Standing wave modes#
Looking at the spatial component of waves, based on the length of the transmission line $X$, standing waves form such that:
$$ X = m\frac{\lambda}{2} \quad $$and
$$ k = \frac{m\pi}{X}, \quad m = 0,1,2,\dots $$where $m$ represents the integer $\pi$ multiples of standing waves, and $k$ is our wavenumber. So, how far apart are these modes (wavenumbers)? We are solving for the spacing between the increments of m $k_{m+1} - k_m$, which is just taking the derivative of the wavenumber with respect to $m$:
$$k_{m+1} - k_m = \frac{d}{dm}\left(\frac{m\pi}{X}\right) $$$$ \frac{d}{dm} = \boxed{\frac{\pi}{X}} \tag{1.2} $$Assuming $X \gg \lambda$ such that we can fit a sufficient amount of modes on our transmission line (think squeezing a bunch of points close together the same way you would approximate the area under a curve with rectangles to lead up to a continuous integral), we can express the number of modes on the transmission line as
$$ \boxed{N_k = \frac{X}{\pi}} \tag{1.3} $$or simply the inverse of the spacing between modes.
Now, recall the relationship between wavelentgh and frequency $\lambda = \frac{c}{f}$. Here, $c$ is the speed of light in Copper, known as phase velocity. We will denote this quantity as $c^\prime$ moving forward so we don’t forget.
2. From Wavenumber to Frequency — Density of States#
Recall $\lambda = c/f$, where $c$ is the phase velocity — here in copper, denoted $c’$.
We want this in terms of frequency:
$$ X = m\frac{\lambda}{2} \Rightarrow \lambda = \frac{2X}{m} $$$$ N_k = \frac{X}{\pi} \Rightarrow X = \pi N_k $$Combining:
$$ \lambda = \frac{2\pi N_k}{m} = \frac{c'}{f} $$$$ \;\Rightarrow\; f = \frac{mc'}{2\pi \cancel{N_k}^{\to \frac{\pi}{X}}} = \frac{\cancel{\pi} m c'}{2\cancel{\pi} X} $$$$ \boxed{f = \frac{mc'}{2X}} \tag{2.1} $$This is a quantized frequency. Find the spacing between these frequencies the same way as the wavenumber spacing (spacing between one mode, with respect to $m$):
$$f_{m+1} - f_m = \frac{d}{dm}\left(\frac{mc'}{2X}\right) = \boxed{\frac{c'}{2X}}$$Again, because $X \gg \lambda$, the number of possible modes is the inverse of the spacing:
$$\boxed{N_f = \frac{2X}{c'}}$$Both of these results are tied to a unit length. To generalize, express the result as a density of states per unit length (as in solid state physics). We do this because densities are invariant when you scale. Since we’re working in one dimension:
$$D(E) = \frac{N_f}{L} = \frac{2\cancel{X}/c'}{\cancel{X}} \;\Rightarrow\; \boxed{D(E) = \frac{2}{c'}}$$Electromagnetic energy is transferred in discrete packets called photons, with photon energy $E = hf$ ($h$ = Planck’s constant). The total energy is comprised of $N$ photons:
$$E_N = \sum_{n=0}^{N} nhf \;\Rightarrow\; \boxed{E_N = Nhf}$$3. Photon Probability Distribution#
With $N$ possible energy states, we need the probability that the system is in a given energy state, as a function of total energy and temperature, $P(E,T)$. Constrain the system to thermal equilibrium — average energy is fixed, and fluctuations occur only along the temperature axis.
Use the Boltzmann–Gibbs distribution:
$$\rho(E_N) \propto e^{-E_N/kT}$$where $k$ = Boltzmann’s constant $= 1.38\times10^{-23}\ \text{J/K}$, and $T$ = temperature in Kelvin.
Plugging in $E_N$:
$$P(E_N) \propto e^{-Nhf/kT}$$This is only a proportionality. The actual probability function is the ratio of this Boltzmann factor (for a single energy state) to the total energy of the system — an infinite sum over all states:
$$E_{TOT} = \sum_{N=0}^{\infty} e^{-Nhf/kT}$$$$P(N) = \frac{E_N}{E_{TOT}} = \frac{e^{-Nhf/kT}}{\displaystyle\sum_{N=0}^{\infty} e^{-Nhf/kT}}$$Aside: the geometric series#
The denominator is a geometric series $\sum y^i$. Recall the general power series:
$$\sum_{i=0}^{\infty} ay^i = a + ay + ay^2 + ay^3 + \cdots$$Here $a=1$, so consider $S = \sum_{i=0}^{\infty} y^i = 1 + y + y^2 + \cdots$
Represent as a finite sum first, then take the limit $N\to\infty$:
$$S = \sum_{i=0}^{N} y^i = y^0 + y^1 + \cdots + y^N$$Multiply both sides by $y$:
$$yS = y + y^2 + \cdots + y^{N+1}$$Asserting the sum converges, $yS = \sum_{i=0}^{N} y^{i+1} \approx \sum_{i=0}^{N} y^i = S$ (not exact — adding one extra term). Back out the exact difference:
$$S - yS = \left(1+y+\cdots+y^N\right) - \left(y+y^2+\cdots+y^{N+1}\right) = 1 - y^{N+1}$$$$S(1-y) = 1-y^{N+1} \;\Rightarrow\; S = \frac{1-y^{N+1}}{1-y}$$Since we’re dealing with an infinite series:
$$S = \sum_{i=0}^{\infty} y^i = \lim_{N\to\infty}\frac{1-y^{N+1}}{1-y}$$For this to converge, we need $|y| < 1$ (otherwise it diverges). With that constraint:
$$\boxed{S = \sum_{i=0}^{\infty} y^i = \frac{1}{1-y}} \qquad (|y|<1)$$Back to the probability function#
$$P(N) = \frac{e^{-Nhf/kT}}{\sum_{N=0}^\infty e^{-Nhf/kT}} = e^{-Nhf/kT}\left(1-e^{-hf/kT}\right)$$$$\boxed{P(N) = e^{-Nhf/kT}\left(1-e^{-hf/kT}\right)}$$This is the Photon Probability Distribution — it describes how likely you are to find a photon at a given energy level.
4. Average Number of Photons per Mode (Bose–Einstein Distribution)#
The next question: what is the average number of photons at a given frequency (energy level)? Multiple photons can occupy the same energy level, and the number of photons per mode is governed by the distribution above. The average is weighted by this distribution:
$$\langle N \rangle = \sum_{N=0}^{\infty} N\,P(N)$$Writing $y = e^{-hf/kT}$, the distribution takes the general form $P(N) = y^N(1-y)$, so:
$$\langle N \rangle = \sum_{N=0}^{\infty} N y^N(1-y) = (1-y)\sum_{N=0}^{\infty} N y^N$$Aside: solving $\sum N y^N$ by differentiating the geometric series#
We already know $\sum_{i=0}^{\infty} y^i = \dfrac{1}{1-y}$ for $|y|<1$. Differentiate both sides w.r.t. $y$:
$$\frac{d}{dy}\sum_{i=0}^{\infty} y^i = \sum_{i=1}^{\infty} i\,y^{i-1}$$(the sum on the right starts at $i=1$ since the $i=0$ term is constant, and its derivative is zero.)
Differentiate the closed form using the chain rule, with $u = 1-y$, $f(u) = 1/u$, $g(y) = 1-y$:
$$\frac{d}{du}(u^{-1}) = -u^{-2} = -\frac{1}{u^2}, \qquad \frac{d}{dy}(1-y) = -1$$$$\frac{d}{dy}\left(\frac{1}{1-y}\right) = \frac{1}{u^2} = \boxed{\frac{1}{(1-y)^2}}$$So now we have a result for $\sum i,y^{i-1}$, but we want $\sum i,y^i$. To go from $y^{i-1}$ to $y^i$, multiply by $y$:
$$y\cdot y^{i-1} = y^i$$$$\sum_{N=0}^{\infty} N y^N = y\sum_{i=1}^{\infty} i\,y^{i-1} = \frac{y}{(1-y)^2}$$Putting it together#
$$\langle N \rangle = (1-y)\cdot\frac{y}{(1-y)^2} = \boxed{\frac{y}{1-y}}$$Substituting back $y = e^{-hf/kT}$:
$$\langle N \rangle = \frac{e^{-hf/kT}}{1-e^{-hf/kT}} = \boxed{\frac{1}{e^{hf/kT}-1}}$$The angle-bracket (ket-like) notation indicates an expectation value, as used in probability distributions and quantum mechanics.
This is the Bose–Einstein distribution. It underlies the Bose–Einstein Condensate (BEC), which is the core mechanism behind many neutral-atom quantum technologies. Its general form is:
$$\langle N \rangle = \frac{1}{e^{(E-\mu)/kT}-1}$$where $E$ is the quantum state energy and $\mu$ is the chemical potential of the species of atom used. When $E = \mu$, $\langle N\rangle \to \infty$ — a macroscopic number of bosons accumulate in the lowest energy state, which is the definition of a BEC.
For the Johnson noise problem, this simplifies because photons have no chemical potential, so $\mu = 0$.
5. Average Energy Density (per Unit Length and Bandwidth)#
Recap — three results derived so far:
- Modes per unit length & bandwidth (density of states): $\quad D(E) = \dfrac{2}{c’}$
- Photons per mode (Bose–Einstein): $\quad \langle N \rangle = \dfrac{1}{e^{hf/kT}-1}$
- Average photon energy: $\quad E = \langle N\rangle hf$
The goal is to combine these into the total energy density per unit length and bandwidth — zooming back out of the quantum picture to the macroscopic one, while keeping generality via density:
$$\langle E \rangle = D(E)\cdot E = \frac{2}{c'}\cdot hf \cdot \frac{1}{e^{hf/kT}-1}$$$$\boxed{\langle E \rangle = \frac{2}{c'}\cdot\frac{hf}{e^{hf/kT}-1}}$$6. From Energy Density to Power (Single Resistor)#
Recall the experimental setup: two resistors, each emitting EM energy. Because the system is in thermal equilibrium, each resistor accounts for exactly half of the average energy.
To convert average energy to average power, do dimensional analysis. Average energy has units:
$$\langle E \rangle = \frac{J}{m\cdot Hz}$$We want power per unit bandwidth, $\bar{P}$, with units $\dfrac{W}{Hz} = \dfrac{J}{s\cdot Hz}$.
Power is a flow of energy per unit time (i.e., speed), and the wave speed is $c’$. Bringing it together (the factor of $\tfrac12$ accounts for a single resistor):
$$\bar{P} = \frac{1}{2}\langle E\rangle c' = \frac{1}{2}\cdot\frac{2}{c'}\cdot\frac{hf}{e^{hf/kT}-1}\cdot c'$$$$\boxed{\bar{P} = \frac{hf}{e^{hf/kT}-1}} \qquad \left[\frac{W}{Hz}\right]$$This is a power density (W/Hz), so define power as a function of frequency:
$$P_f = \bar{P}\,\Delta f = \frac{hf}{e^{hf/kT}-1}\,\Delta f$$Here $f$ acts as an effective center frequency — a coarse knob indicating which part of the EM spectrum you’re in (e.g., microwave vs. optical) — and $\Delta f$ is the bandwidth window around it. For the power distribution to be uniform across this window, we need $\Delta f \ll f$.
Asserting this uniformity means $P_f = \bar{P}\Delta f$ remains a good approximation, so we don’t need to evaluate the full integral:
$$P_f = \int_f^{f+\Delta f} \bar{P}(f')\,df'$$When comparing the frequency spectrum of electronics to the entire EM spectrum, we’re practically at DC — even in the microwave realm. This means $hf \ll kT$ in the exponential’s denominator, and $\Delta f$ can be treated as a small signal about an effective Q-point, just like in semiconductor devices / DC operating points of DC-DC converters.
7. Taylor Series (Rayleigh–Jeans Limit)#
Perform a first-order Taylor expansion of the exponential about the Q-point:
$$y(a) = f(a) + \frac{f'(a)}{1!}(x-a)$$- 0th order: $e^0 = 1 \Rightarrow$ coefficient $= 1$
- 1st order: $f’(a) = f’(e^x) = e^x \Rightarrow f’(0) = e^0 = 1 \Rightarrow$ coefficient $=1$
Thus:
$$\boxed{e^{hf/kT} \approx 1 + \frac{hf}{kT}}$$This is the Rayleigh–Jeans limit — the first classical approximation attempting to explain blackbody radiation. Historically, this approximation first appeared in the 3-D version of this exact problem, and it led to the ultraviolet catastrophe, where the model predicted infinite energy at ultraviolet frequencies and higher. This breakdown led to the birth of quantum mechanics, with Planck, Bose, and Einstein providing the fix. The Bose–Einstein distribution ultimately resolved the catastrophe, but the key insight was Planck’s realization that energy is quantized into discrete packets — photons.
Plugging back into $P_f$#
$$P_f = \frac{hf}{e^{hf/kT}-1}\,\Delta f = \frac{hf}{1+\frac{hf}{kT}-1}\,\Delta f = \frac{hf}{hf/kT}\,\Delta f$$$$\boxed{P_f = kT\,\Delta f}$$This is the average power per mode emitted by a single resistor.
8. Relating Power to Voltage — the Johnson Noise Result#
Relate $P_f$ to the traveling wave amplitude $V_0$. Since current flows through both resistors: $V_0 = I_0 \cdot 2R$.
$$I_0 = \frac{V_0}{2R} \;\Rightarrow\; P_f = \left(\frac{V_0}{2R}\right)^2 R = \frac{V_0^2}{4R^2}\cdot R$$$$\boxed{P_f = \frac{V_0^2}{4R}}$$$V_0$ here is the Johnson noise mean-square voltage, henceforth denoted $V_n^2$.
Setting the two expressions for $P_f$ equal:
$$\bar{P} = kT\Delta f = \frac{V_n^2}{4R}$$Solving for the Johnson noise voltage (of a single resistor):
$$\boxed{V_n = \sqrt{4kTR\Delta f}}$$Noise density form#
$\Delta f$ is the equivalent bandwidth of concern for a given circuit, so it’s most convenient to express this as a noise density, leaving out $\sqrt{\Delta f}$ (or $\sqrt{BW}$):
$$\boxed{e_n = \sqrt{4kTR}}$$where $T$ = temperature in Kelvin ($300,K$ at room temperature), $k$ = Boltzmann’s constant $= 1.38\times10^{-23}\ \text{J/K}$.
Equivalently, as a current noise density:
$$\boxed{i_n = \sqrt{\frac{4kT}{R}}}$$Units: $e_n$ in $V/\sqrt{Hz}$, $i_n$ in $A/\sqrt{Hz}$.
If this derivation doesn’t give you a true appreciation for how much human advancement was needed to quantify the noise in a single resistor, then you’re not recognizing the art here. Quantum physics had to be conceptualized before we could arrive at this beautifully simple result.
9. Circuit Models#
- Thevenin equivalent: voltage noise source $e_n$ in series with resistor $R$.
- Norton equivalent: current noise source $i_n$ in parallel with resistor $R$.
10. Looking Ahead: Shot Noise#
Shot noise was first observed in vacuum tubes by Walter Schottky, who developed the theorem describing this phenomenon. It is a quantum effect present in semiconductor devices due to the discrete nature of electrons.
DC current through a P-N junction experiences random fluctuations because minority carrier density is constantly changing due to thermal effects and random generation/recombination of charge carriers. This is a Poisson process — truly stochastic (random). Each electron arrives at a random time and can be described as a single impulse response of charge $q$. All of these impulses sum together to give a total current with respect to time, $I(t)$:
(derivation continues beyond the transcribed notes)

