ENZH
Discuss this post with AI
ChatGPTClaude

What's Actually Inside an Optical Module

I recently listened to an episode of the Chinese podcast 十分吸引 on optical communication. The guest, Kevin, is a chip systems engineer whose PhD was in analog circuit design. He walks from the fiber box on your living-room wall all the way to co-packaged optics, and by the end you understand why the network inside a data center has become part of the compute. I look at GPU cluster bills for a living. The part of the bill that isn't the GPU, I knew only roughly — switches, and a lot of optical modules — and I had never really asked what each of those modules is doing. Listening, it clicked that this is the same shape as the memory wall I wrote about earlier: the GPU is waiting on data, and whether it's waiting on memory or on a card in the next rack, it's waiting. So I turned the episode into a series. This is part one, and the job here is just to say what an optical module is.

You already own one

Optical modules get a lot of press, and the word makes them sound like something that lives in a hyperscaler's basement. Every household with fiber has one. The small box on the wall with a thin fiber going into it and a light blinking — the ONT, or "fiber modem" — is an optical module, just a slow one. Kevin's point on the show is that it's a few dozen times slower than the high-speed modules in a data center, but the basic construction inside is the same.

Inside it is a single-fiber bidirectional optical assembly. On the transmit side there's a laser that turns your outgoing bits into light. On the receive side there's a photodetector, a PD, that turns incoming light back into electricity. Next to the PD sits an amplifier stage, a transimpedance amplifier or TIA, because the light that arrives is faint and the current the PD produces is in the microamp range — nothing downstream can work with it until it's been amplified. Send and receive share the one fiber by using two different wavelengths, one upstream and one downstream, so they don't collide.

That fiber leaves your apartment and runs to a passive splitter somewhere in the building or neighborhood. Passive means it has no power supply; its only job is to merge a few dozen households onto one fiber, typically at a 1:32 or 1:64 ratio, and that one fiber runs to the operator's central office. From the central office to your wall it is light the whole way.

For scale: China had 691 million fixed broadband subscribers at the end of 2025, and by mid-2025 fiber-to-the-home ports made up 96.6% of all fixed broadband ports. That's the largest optical access market anywhere. The technology has moved through generations too: early GPON offered 2.5G of total downstream bandwidth shared across a few dozen homes; 10G PON is what made gigabit-to-the-home possible; and China Mobile ran 50G PON trials in 2025, with 10-gigabit residential service on the way.

Dozens of households take turns on one fiber, which is why a 10-gigabit tier is a burst peak rather than a guaranteed slice — nobody saturates their link at the same time, so offering 10 gigabit per home under a 50G PON works, and if everyone actually went full-rate at once there wouldn't be enough to go around.

孙悦 brought up dialing in during the 1990s, ADSL, ISDN, all of it running over phone lines, all of it electrical. The show picked that up: the first transatlantic telephone cable, TAT-1, went into service in 1956 with 36 channels, and by the seventies and eighties submarine cables carried thousands of simultaneous calls — but those early cables were electrical too: on both engineering difficulty and cost, electricity is always the one that arrives first. Fiber initially reached only the key nodes of the national backbone — Shanghai would have a backbone interface, say — while every cell tower and every home below it stayed on copper. Once users' bandwidth demands rose, everything eventually had to go optical, because fiber's loss per kilometer is tiny and fixed, so total attenuation climbs only slowly as the run gets longer. The reason older neighborhoods are slow to get fiber is the same ledger: trenching and laying cable costs money, and whether the operator bothers depends on who's paying. 石磊 compared it to retrofitting elevators onto old apartment blocks — some blocks push it through, some never do.

There was a picture on the show I found useful: the whole optical network is a circulatory system, and the fiber box on your wall is the capillary at the far end. So what's the aorta, and how big is it?

The optical modem on your wall is the capillary at the very edge of the optical network; the modules packed into leaf and spine switch panels in a data center are the aortaThe optical modem on your wall is the capillary at the very edge of the optical network; the modules packed into leaf and spine switch panels in a data center are the aorta

The aorta is in the data center

Start with a conventional cloud data center. A server connects first to the switch at the top of its rack. That run is usually under three meters and is still copper — a direct attach cable, or DAC. From that top-of-rack switch upward, though, there's a leaf layer and a spine layer of switching, and between those layers it's all light, with 100G and 400G modules filling every faceplate. The traffic flowing sideways between servers, what the industry calls east-west traffic, is generally put at 70 to 80 percent or more of everything inside a data center, and in AI clusters it's higher still.

AI clusters push this much further. Training a large model means wiring tens of thousands of GPUs into a single network — the scale-out network, in the trade's vocabulary. GPUs compute in parallel, and the show's image was a row of wide-open mouths all eating at once; they are extraordinarily hungry for bandwidth, for the same reason HBM has to be stacked right next to the GPU — on-package and off-package, everything is fighting for bandwidth. A few years ago a GPU didn't need many optical modules, because all the cards sat inside one rack, centimeters apart, and there was nothing to bridge. Now it's a hundred thousand cards. A rack holds a few dozen — an NVL72 rack is 72 — so a hundred thousand cards runs past a thousand racks, which fills an entire floor. A server in one corner of that floor needs to talk to a server diagonally across the room, tens or hundreds of meters away, and the latency has to stay low or the cards can't cooperate on one training job. That link can only be light.

High bandwidth and many tiers of switching together produce the number Kevin gave on the show: on average, one GPU pulls along two to three, sometimes more, 800G or 1.6T optical modules. The exact count depends heavily on topology — a two-tier network lands around 1:1, a three-tier leaf-spine can reach 1:6 — but the direction is right, and the consequence is that optical-module demand now hangs directly off GPU shipments.

Only some data centers need modules this expensive. 孙悦 asked whether a facility that's huge and full of servers but isn't training frontier models still needs all that long-reach communication. 敏姐's answer was no: something as small as a company's own server room, a handful of machines, is nearly all copper; a storage-heavy facility still needs optical modules, just not fast ones. 800G and 1.6T are specifically what AI training and inference dragged into existence.

Opening up an 800G

Now take a real 800G module apart. Kevin actually brought an old multimode module along on the show: pry off the black dust cap, one end takes a fiber, the other end is a gold-finger edge connector that plugs into the switch faceplate. Light in one end, electricity out the other. In Kevin's words, what an optical module does is optical-to-electrical conversion for communication, and that's the whole job.

Inside is a symmetric signal chain. Take the transmit direction first, electricity becoming light. Eight lanes of 100G-class PAM4 electrical signal arrive at the gold fingers at the tail of the module — 106.25 gigabits per lane once you count the FEC overhead, which is why the industry calls this a 112G-class SerDes. PAM4 is pulse-amplitude modulation with four voltage levels — each level encodes two bits, 00, 01, 10, 11 — so the symbol rate on each lane is half the bit rate. Those signals have already traveled a few dozen centimeters across the host PCB. Kevin's image for that board is a road, and it is a bumpy one — by the time the signal reaches the connector the waveform is badly smeared. So the first stop is a DSP, a digital signal processing chip, which equalizes and retimes the signal to clean it up. Next is a driver, which amplifies the clean but small signal into a swing large enough to push a laser. Then comes the laser or modulator itself — an EML, an electro-absorption modulated laser, or a silicon photonic modulator — and that's where electricity becomes light. Finally the light is coupled into the fiber through the transmit optical subassembly, the TOSA, and it's on its way.

The receive direction retraces the path: the PD turns light back into current, microamps again; the TIA right behind it amplifies that to a usable voltage; then the DSP's ADC samples, equalizes and makes bit decisions, and the result leaves through the gold fingers.

There's also an MCU, a small management chip acting as the module's physician, reporting temperature, optical power and laser bias current in real time. If the laser is an EML there's usually a thermoelectric cooler, a TEC, holding the wavelength steady, though not every part number has one.

Open an 800G optical module and you find one symmetric signal chain: electrons in at the gold fingers, cleaned by the DSP, turned into light by the laser; on the way back the detector turns light into current and the DSP hands it out againOpen an 800G optical module and you find one symmetric signal chain: electrons in at the gold fingers, cleaned by the DSP, turned into light by the laser; on the way back the detector turns light into current and the DSP hands it out again

This chain isn't always four separate chips anymore. The newer generation of PAM4 DSPs has absorbed the TIA and the laser driver into the DSP die. On power, an 800G module draws around 16 watts, and the DSP alone is close to half of that — the single largest consumer inside. That comes back in the packaging war, because the electricity that co-packaged optics is trying to save is exactly this.

Reading a part number

Reading a module's part number is two steps. First the package, the physical form. For 800G the two main "cages" are QSFP-DD and OSFP. QSFP-DD is 18.35 mm wide and has the larger installed base, because it's backward compatible with 200G and 400G. OSFP is 22.58 mm wide, slightly larger, with more surface for heat dissipation, so high-power AI deployments prefer it, and 1.6T is heading almost entirely to OSFP — QSFP-DD1600 exists and is more backward-compatible, but it can't shed 1.6T's heat.

Second the suffix, which is the optical scheme. SR is short-reach multimode, 850 nm, VCSEL lasers, under a hundred meters. DR is parallel single-mode, 500 meters. FR is 2 km, LR is 10 km, and those three all use the 1310 nm O-band. ZR is coherent, 80 km and beyond, on the 1550 nm C-band. The digit in the suffix is the place to be careful, because it is not always a lane count. DR8 and SR8 really are eight parallel lanes, 8 × 100G on eight fiber pairs. FR4 and LR4, though, are four lanes at 200G each, and those four are wavelength-multiplexed onto one fiber pair. Parallel fibers and wavelength-division multiplexing are two different ways to get to 800G, and they get confused constantly. Put it together and "800G DR8" reads as 800G total, eight lanes, 100G per lane, eight parallel single-mode pairs, 500 meters — the workhorse for connecting things on the same floor of a data center.

So is 1.6T just twice the lane count? 孙悦 asked exactly that. 敏姐's answer: 1.6T is not 16 × 100G, it's 8 × 200G. The reason is mundane. A faceplate only fits so many fibers, one ribbon can integrate about eight of them, and the switch faceplate's port openings have physical dimensions too. So the only route is to carry double the bandwidth over the same number of fibers, which means doubling the rate per lane. And when the lane rate doubles, every chip on both ends — DSP, driver, laser, TIA — has to double its speed and bandwidth, so each new generation of module means redesigning basically every chip inside it. 敏姐's comment was that the analog design difficulty here climbs exponentially.

Price, briefly. A third-party 800G DR8 runs roughly $900 to $1,200; OEM-branded ones can go past $6,000. An 800G passive copper DAC of half a meter to two meters is $105 to $300. Even comparing the cheapest module to the priciest cable, the module costs several times more; at the branded end it is more than an order of magnitude.

So why can't electricity go far?

孙悦's summary at the end of this stretch of the show was the right one: electricity doesn't travel far, light does, so every time data has to go a distance, electricity gets translated into light.

Which leaves a question hanging. Why doesn't electricity travel far? Copper is fine for three meters, and by a few dozen meters you have to switch to light — where exactly does the wire hit its limit? That's the next post.


This series is compiled from episode 75 of the Chinese podcast 十分吸引, "光与电的游戏:有线通讯史的百年之争", with guest Kevin, a chip systems engineer, and hosts 石磊, 敏姐 and 孙悦. The framing is Kevin's; I reorganized it by theme and wrote it up, so any errors are mine. Neither the episode nor this post is investment advice — companies are named only as examples of where the supply chain sits.

Discuss this post with AI
ChatGPTClaude
Light vs CopperPart 1 of 5
← PrevNext →

Subscribe

New posts straight to your inbox. Nothing else.


© Xingfan Xia 2024 - 2026 · CC BY-NC 4.0