Two chapter-7 devices, drains tied together
Take one NMOS and one PMOS, both the square-law devices from chapter 7. Tie the two gates together — that is the input, Vin. Tie the two drains together — that is the output, Vout. Put the NMOS source at ground and the PMOS source at VDD. That is the whole circuit. It is called a CMOS inverter, and it is the single most repeated structure on any digital chip: every logic gate, every flip-flop, every SRAM cell is built from a handful of these wired together.
Read it as a pair of switches stacked between the rails. When Vin is near 0 V, the NMOS is off (VGS below its threshold) and the PMOS is hard on (its VSG is nearly the whole rail) — the output is pulled to VDD. When Vin is near VDD, the roles swap and the output is pulled to ground. In between, both devices are partway on at once, and that narrow region is where all the interesting physics of this chapter lives.
The point where the two switches trade places — where the transfer curve crosses Vout = Vin — is called the switching threshold, VM. It is not automatically at VDD/2: it lands wherever the NMOS's pull-down strength balances the PMOS's pull-up strength, and nothing so far has said those are equal.
Commit before you touch anything
Hold Vin at exactly VM, the point where the transfer curve crosses the diagonal. What state are the two transistors in?
Why the crossing isn't automatically in the middle
Silicon does not give holes and electrons the same mobility — holes move at roughly 40% of the speed electrons do in the same field. A PMOS built with the same width and length as its NMOS partner is a substantially weaker pull-up. Left uncorrected, that drags VM up toward VDD: the NMOS wins the tug of war earlier, because it barely has to try. The standard fix is exactly the one real layouts use — draw the PMOS wider, by roughly the mobility ratio, until the two pull strengths match. It is why, if you have ever looked at a standard-cell layout, the PMOS row is visibly fatter than the NMOS row.
In the wild
The 7404 hex inverter, and every NOT gate since. Six of exactly this circuit in one package, unchanged in principle since the first CMOS logic families of the 1970s.
Standard-cell layout. Open any cell library's inverter and the PMOS transistor is drawn visibly wider than the NMOS — the mobility correction from this lesson, done in silicon rather than in a slider.
Schmitt-trigger inputs on microcontroller pins. Built from deliberately missized inverters, so the switching threshold on the way up differs from the way down — hysteresis, purchased with the same Wn/Wp knob this bench turns.
Before you move on
You make the PMOS much wider without touching the NMOS. What happens to VM?
No resistors, and (almost) no current
Look again at the topology: NMOS and PMOS, drains tied together, no resistor anywhere. That is not an oversight — it is the entire reason CMOS displaced every logic family before it. At either rail, one device is hard on and the other is hard off, so the series path from VDD to ground is broken by a device sitting in cutoff. No current flows through an open switch, no matter how hard the closed one is pulling.
Compare that with chapter 6's push-pull stage, which needed a small standing bias current through both output devices just to avoid crossover distortion, or with chapter 5's resistor-biased common-emitter stage, which draws current from the supply continuously, whether or not anything is happening at the input. A CMOS gate sitting at either logic level draws nothing from its supply beyond the tiny leakage that lesson 5 takes up on its own.
The exception is the region this chapter has already found interesting for another reason: near VM, both devices are on at once, in saturation, in series. For a narrow slice of Vin around the crossing, there is a real path from rail to rail, and real current flows through it — drawn straight from the supply and dumped straight to ground, doing no useful work at all.
Commit before you touch anything
Sweep Vin slowly from 0 to VDD and watch the current drawn from the supply. What does the curve look like?
Series, not parallel
Because there is no third path out of the shared drain node at DC, the NMOS's current and the PMOS's current are not independent quantities that happen to be measured separately — they are the same current, forced equal by Kirchhoff's current law the instant you tie the drains together. The bench's two current traces do not merely track each other closely; they sit exactly on top of one another, for the same reason chapter 10's two nearly-equal CMRR traces did not need rescaling to look convincing — here it is not two large numbers that happen to be close, it is one number measured twice.
In the wild
Input buffer chains. A slow signal arriving at a chip's pin is deliberately passed through a chain of inverters before it reaches logic, specifically to shrink the time spent near VM and cut this spike down.
Bus contention. Two drivers fighting over one wire is this same rail-to-rail path, but external rather than internal — and it can draw enough current to damage a pad if it persists.
Why "don't leave inputs floating" is CMOS gospel. An undriven gate input can sit near VM indefinitely, drawing this current the whole time — the reason unused CMOS inputs are always tied high, low, or through a defined resistor, never left open.
Before you move on
Why does the NMOS current trace sit exactly on top of the PMOS current trace on this bench?
The energy per flip does not care how big you built the transistors
Every output node in a real chip drives something — the next gate's input, a length of wire, both. Model that as a capacitor, CL, sitting on the shared drain node. Every time the inverter flips, it either charges that capacitor from 0 to VDD through the PMOS, or discharges it from VDD to 0 through the NMOS. That charging and discharging is where a CMOS chip's dynamic power comes from — and where the widespread instinct that follows from lesson 1 and lesson 2 turns out to be wrong.
The instinct: a bigger transistor pushes more current, so surely it takes more energy to do its job. Chapter 7 already told you gm rises with overdrive and width — more current sounds like more everything. But energy is not current; it is current integrated over the time the job takes, and a bigger transistor finishes the job faster, in almost exact proportion. Whether those two effects cancel exactly is a question with a definite answer, and it is worth deriving before you look at the bench.
Take the charging edge. The charge delivered by the supply to raise the node from 0 to VDD is fixed by the capacitor alone — Q = CLVDD, regardless of the path that charge took to get there. The supply sits at a constant VDD throughout, so the energy it delivers is Q·VDD = CLVDD2. Nowhere in that derivation does the PMOS's width, its k, or its instantaneous current appear. They decide how long the charging takes. They cannot touch how much charge a fixed capacitor needs to reach a fixed voltage.
Commit before you touch anything — this is the chapter's misconception
You double every transistor's width, changing nothing else. What happens to the energy the supply delivers each time the output charges from 0 to VDD?
Where the other half goes
The capacitor only ends up holding ½CLVDD2 of that energy — the rest is dissipated as heat in the PMOS while it charges. Then, on the next edge, the NMOS discharges the capacitor to ground and dissipates that stored ½CLVDD2 as heat too, delivering nothing back to the supply. A full 0→1→0 cycle therefore dissipates a full CLVDD2, split evenly between the two transistors, and none of it depends on which transistor did the dissipating or how wide it was built. Multiply by how often a node switches per second and you get the familiar figure, Pdynamic = α·CLVDD2·f, where α is the fraction of clock cycles that actually see a transition.
In the wild
Why "just make it bigger" is free, up to a point. Chip designers routinely upsize a critical gate to hit a timing target with no dynamic-power penalty — the reason timing closure and power closure are separate, largely independent jobs.
Dynamic voltage and frequency scaling (DVFS). Because Pdynamic goes as VDD2, dropping the supply a little saves far more power than dropping the clock the same fraction — the whole reason phones throttle voltage first.
Clock gating. Since α sits right there in the formula, stopping the clock to an idle block multiplies its dynamic power by zero — the single most effective power-saving trick in a modern SoC.
Before you move on
A designer replaces a gate's transistors with ones 4× wider to hit a timing deadline. What happens to that gate's dynamic energy per transition?
What width actually buys
Lesson 3 showed what width does not buy. Here is what it does. During the discharge edge, the NMOS pulls the output from VDD toward ground by supplying current IN(Vout) to a capacitor of size CL. The rate of discharge is dVout/dt = −IN/CL. Since IN scales with W (chapter 7's square law, k ∝ W/L), doubling the NMOS's width roughly doubles the current available at every point on that curve, and the time to cross any fixed voltage span — conventionally, from VDD down to VDD/2, the propagation delay tpHL — falls by roughly the same factor.
Chapter 7's other result matters here too. So long as the discharge stays in saturation the whole way to VDD/2 — VOV comfortably below half the rail — the current is nearly constant across the transition, and the RC intuition of "resistor discharging a capacitor" is only a rough analogy: the actual device is a current source, not a resistor, until it drops into triode near the very end. If the swing crossed into triode earlier, the same voltage-controlled-resistor behaviour from chapter 7's lesson 6 would take over and the delay calculation would look more like an actual RC time constant. This bench's default sizing keeps the whole VDD→VDD/2 swing in saturation, so it behaves like the constant-current case.
Commit before you touch anything
You quadruple the NMOS width only, leaving the PMOS and CL unchanged. What happens to tpHL — the fall-time delay?
In the wild
Buffer chains driving long wires or big loads. A single gate cannot efficiently drive a large off-chip capacitance directly; a tapered chain of progressively wider inverters, each roughly 3–4× the last, gets there faster and at lower total energy than one giant gate.
Clock tree buffers. The literal biggest transistors on most digital chips exist for exactly this reason — driving a clock signal to thousands of destinations with minimum delay skew.
Why a "weak pull-up" I2C bus is slow. Deliberately undersized pull-up strength trades this chapter's speed for lower idle power — the same width knob, turned the other way on purpose.
Before you move on
Which of these speeds up a gate's propagation delay without changing its dynamic energy per transition at all?
The other curve
Everything so far has assumed a device below threshold conducts nothing. Real devices conduct a little even there — an exponentially small but never-zero subthreshold current, the same exponential shape as a diode's reverse leakage, governed by how many VT-widths of margin sit between VGS and VTH. Every "off" transistor on a chip leaks a little, all the time, whether or not the chip is doing anything — this is static power, and it is a genuinely separate curve from lesson 3's dynamic power, related to it only by both being called "power."
The two curves depend on almost disjoint sets of things. Dynamic power cares about CL, VDD2, and how often the gate switches — it vanishes at zero frequency. Static power cares about VTH and VDD, and it does not care about frequency at all: an idle chip still burns it, every second it is powered, doing nothing. The two curves cross at some frequency, above which dynamic power dominates and below which leakage does — and that crossing point is exactly where the low-power design conversation actually happens.
The lever that makes static power interesting is the same VTH this chapter has been sizing around since lesson 1. Lowering VTH gives every gate more overdrive at the same VDD, which lesson 4 says buys speed. But subthreshold current falls off exponentially in (VGS − VTH), so a lower VTH also means the "off" device is closer to "on," and leakage rises exponentially to match. There is no way to buy the first without paying the second.
Commit before you touch anything
You lower VTH by roughly 170 mV to speed up every gate on a chip, holding VDD and frequency fixed. What happens to the crossover frequency where static power equals dynamic power?
In the wild
Multi-threshold CMOS (HVT/SVT/LVT libraries). Real chips mix several threshold voltages on one die — low-VTH cells only on the timing-critical paths, high-VTH everywhere else — buying lesson 4's speed only where it is worth this lesson's leakage cost.
Power gating and sleep transistors. Cutting an idle block's supply rail entirely is the only way to beat subthreshold leakage outright, since lowering VDD alone still leaves a leakage path.
Why your phone's SoC has an idle power spec at all. That number is a direct read of the static curve, measured with the dynamic term forced to zero by clock gating.
Before you move on
A chip is running at a very high clock frequency, flat out. Which power term dominates?
What CMOS costs you that bipolar didn't
Interlude II introduced σ(ΔVth) = AVT/√(WL) and gave it three real, dated numbers — 4.8, 2.4 and 0.9 mV·µm for 0.18 µm, 45 nm and 22 nm process classes — but deliberately never exercised the formula on an actual device. This bench is that device. Every inverter you have built in this chapter is a pair of MOSFETs with some threshold voltage; Pelgrom's law says two of them, built side by side on the same die, will never have exactly the same one. The random part of that mismatch shrinks with the square root of gate area, exactly as it did for the resistor-versus-transistor comparison in the interlude — only now it is VTH itself doing the varying, and VM that inherits the spread.
A single inverter used as a logic gate does not care. Its output only has to land unambiguously on one side of the next gate's own switching threshold; a few millivolts of VM spread across a billion transistors is invisible to a 1-or-0 decision. That is exactly why logic transistors are built at or near a process's minimum size — the smallest area, and therefore the worst absolute matching the process can produce, and nobody has ever needed to fix it.
It becomes a real problem the moment a CMOS pair is asked to do something analogue with that threshold: a self-biased inverter used as a reference, a sense amplifier deciding which bit line is higher, the CMOS mirror chapter 9 already warned you about. There, VM spread is not cosmetic — it is the whole error budget, and the only lever that shrinks it is the one Interlude II priced: area.
Commit before you touch anything
You keep the same W/L ratio (so drive strength and VM are unchanged) but scale both W and L up together, growing the device's area. What happens to the piece-to-piece spread in VM across many copies of this inverter?
Three debts, paid
To chapter 9: a MOS current mirror has no base-current error at all — a MOSFET's gate draws no DC current, so the exact 1/(1+2/β) correction chapter 9 built an entire lesson around simply does not exist here. But what replaces it is worse in a different currency: current mismatch between two MOS mirror legs runs as ΔI/I ≈ 2ΔVTH/VOV, and this bench's own numbers show why that stings — a few millivolts of VTH spread sits directly against an overdrive of only a few hundred millivolts, where a BJT's exponential law effectively divided its own mismatch by VT, a much larger number in the denominator. No base current, and a worse matching problem: both true, for the same underlying reason.
To chapter 9, again — the rail. A saturated MOSFET needs VDS ≥ VOV just to stay in the region this bench has depended on all chapter. At this bench's default sizing that is roughly 0.48 V of headroom per stacked device, on a 1.8 V rail. Stack a cascode — two devices in series doing one job — and you have spent over half the rail on overdrive alone, before a single other stage gets a volt to work with. Bipolar's VCE,sat of a few hundred millivolts was cheap by comparison; that is the "rail too low to stack cascodes freely" chapter 9 promised you.
To chapter 10: the same long-tailed pair, rebuilt from this chapter's devices, has a tail and a load that are both MOS — and chapter 10's CMRR numbers do not carry over, because they leaned on VA/VT ≈ 2,900. The MOS equivalent, gmro, is this bench's own lesson-1 gain figure: about −52 at VM, two orders of magnitude smaller. Every dB of CMRR chapter 10 built from that ceiling shrinks by the same two orders of magnitude the moment the devices are MOS instead of bipolar.
In the wild
SRAM bit cells. Built at the smallest area a process allows, for density — and threshold mismatch between the cell's own transistors is a leading cause of read/write failures at the tails of a large memory array, which is why SRAM cells are sized slightly above absolute minimum despite the area cost.
Comparators and sense amplifiers. Reach for exactly the enlarged, Pelgrom-priced devices this lesson describes, right next to logic built at minimum size two transistors away on the same die.
Why analogue and digital blocks look so different on a die photo. Interlude II already showed this from the cost side; this lesson is the physical reason the photo looks that way at all.
Before you move on
Two CMOS process classes are compared: one with AVT = 2.4 mV·µm, one with AVT = 0.9 mV·µm. At the same physical device area, which has the tighter VTH match?
Checkpoint
1. At the switching threshold VM, a well-designed inverter's two transistors are…
2. A CMOS gate sitting steady at either logic level draws, at DC (ignoring leakage)…
3. Doubling every transistor's width on a chip, everything else fixed, roughly…
4. Lowering VTH to speed a chip up, at fixed VDD and frequency, moves the dynamic/static crossover frequency…
5. ch 9 A MOS current mirror, compared with chapter 9's bipolar one, has…
6. ch 10 Why don't chapter 10's bipolar CMRR figures carry over to the same pair built in CMOS?