In‑depth analysis of microcontroller security weaknesses, invasive and non‑invasive attack techniques, and practical countermeasures based on real‑world reverse engineering experience.
We describe techniques for extracting protected software and data from smartcard processors, including manual microprobing, laser cutting, focused ion‑beam manipulation, glitch attacks, and power analysis. Many of these methods have already been used to compromise widely‑fielded conditional‑access systems, and current smartcards offer little protection against them. We give examples of low‑cost protection concepts that make such attacks considerably more difficult. A thorough grasp of reverse engineering is essential before any defensive strategy can be properly designed. While no single countermeasure is foolproof, a combination of architectural and physical tricks can make unauthorised access substantially more time‑consuming and expensive.
Smartcard piracy has become a common occurrence. Since around 1994, almost every type of smartcard processor used in European, and later also American and Asian, pay‑TV conditional‑access systems has been successfully reverse engineered. Compromised secrets have been sold in the form of illicit clone cards that decrypt TV channels without revenue for the broadcaster. The industry has had to update the security processor technology several times already and the race is far from over. Smartcards promise numerous security benefits. They can participate in cryptographic protocols, and unlike magnetic stripe cards, the stored data can be protected against unauthorised access. However, the strength of this protection seems to be frequently overestimated. A single successful hack can undermine an entire deployment if the attacker gains persistent control over the key material. In Section 2, we give a brief overview on the most important hardware techniques for breaking into smartcards. We aim to help software engineers without a background in modern VLSI test techniques in getting a realistic impression of how physical tampering works and what it costs. Based on our observations of what makes these attacks particularly easy, in Section 3 we discuss various ideas for countermeasures. Some of these we believe to be new, while others have already been implemented in products but are either not widely used or have design flaws that have allowed us to circumvent them.
We can distinguish four major attack categories:
All microprobing techniques are invasive attacks. They require hours or weeks in a specialised laboratory and destroy the packaging. The other three are non‑invasive attacks. After preparation, they can be reproduced within seconds on another card of the same type without physical harm. Non‑invasive attacks are particularly dangerous because the owner may not notice that secrets have been stolen, and the equipment can be disguised as a normal reader.
Depackaging: Invasive attacks start with the removal of the chip package. We heat the card plastic until it becomes flexible. This softens the glue and the chip module can then be removed easily by bending the card. We cover the chip module with 20‑50 ml of fuming nitric acid heated to around 60 °C and wait for the black epoxy resin that encapsulates the silicon die to completely dissolve. The procedure should preferably be carried out under very dry conditions, as the presence of water could corrode exposed aluminium interconnects. The chip is then washed with acetone in an ultrasonic bath, followed optionally by a short bath in deionized water and isopropanol. We remove the remaining bonding wires with tweezers, glue the die into a test package, and bond its pads manually to the pins. Attempting to crack a modern chip without proper chemical handling often results in destroying the very traces you intend to read.
Layout Reconstruction: The next step in an invasive attack on a new processor is to create a map of it. We use an optical microscope with a CCD camera to produce several meter large mosaics of high‑resolution photographs of the chip surface. Basic architectural structures, such as data and address bus lines, can be identified quite quickly by studying connectivity patterns and by tracing metal lines that cross clearly visible module boundaries (ROM, RAM, EEPROM, ALU, instruction decoder, etc.). All processing modules are usually connected to the main bus via easily recognisable latches and bus drivers. The attacker obviously has to be well familiar with CMOS VLSI design techniques and microcontroller architectures, but the necessary knowledge is easily available from numerous textbooks. Photographs of the chip surface show the top metal layer, which is not transparent and therefore obscures the view on many structures below. Unless the oxide layers have been planarised, lower layers can still be recognised through the height variations that they cause in the covering layers. Deeper layers can only be recognised in a second series of photographs after the metal layers have been stripped off, which we achieve by submerging the chip for a few seconds in hydrofluoric acid (HF) in an ultrasonic bath. HF quickly dissolves the silicon oxide around the metal tracks and detaches them from the chip surface. HF is an extremely dangerous substance and safety precautions have to be followed carefully when handling it. A dedicated adversary can use reverse engineering to recover the entire floorplan within days, given the right imaging tools. Figure 3 demonstrates an optical layout reconstruction of a NAND gate followed by an inverter. These images were taken with a confocal microscope, which assigns different colours to different focal planes (e.g., metal=blue, polysilicon=green) and thus preserves depth information. Multilayer images like those can be read with some experience almost as easily as circuit diagrams. These photographs help us in understanding those parts of the circuitry that are relevant for the planned attack. If the processor has a commonly accessible standard architecture, then we have to reconstruct the layout only until we have identified those bus lines and functional modules that we have to manipulate to access all memory values. More recently, designers of conditional‑access smartcards have started to add proprietary cryptographic hardware functions that forced the attackers to reconstruct more complex circuitry involving several thousand transistors before the system was fully compromised. However, the use of standard‑cell ASIC designs allows us to easily identify logic gates from their diffusion area layout, which makes the task significantly easier than the reconstruction of a transistor‑level netlist. Before you can dump any meaningful data, you must first understand how the address decoder maps physical rows to logical addresses. Some manufacturers use non‑standard instruction sets and bus‑scrambling techniques in their security processors. In this case, the entire path from the EEPROM memory cells to the instruction decoder and ALU has to be examined carefully before a successful disassembly of extracted machine code becomes possible. However, the attempts of bus scrambling that we encountered so far in smartcard processors were mostly only simple permutations of lines that can be spotted easily. Any good microscope can be used in optical VLSI layout reconstruction, but confocal microscopes have a number of properties that make them particularly suited for this task. While normal microscopes produce a blurred image of any plane that is out of focus, in confocal scanning optical microscopes, everything outside the focal plane just becomes dark. Confocal microscopes also provide better resolution and contrast. A chromatic lens in the system can make the location of the focal plane wavelength dependent, such that under white light different layers of the chip will appear simultaneously, but in different colours. Automatic layout reconstruction has been demonstrated with scanning electron microscopy. We consider confocal microscopy to be an attractive alternative, because we do not need a vacuum environment, the depth information is preserved, and the option of oil immersion allows the hiding of unevenly removed oxide layers. With UV microscopy, even chip structures down to 0.1 µm can be resolved. With semiautomatic image‑processing methods, significant portions of a processor can be reverse engineered within a few days. The resulting polygon data can then be used to automatically generate transistor and gate‑level netlists for circuit simulations. Optical reconstruction techniques can also be used to read ROM directly. The ROM bit pattern is stored in the diffusion layer, which leaves hardly any optical indication of the data on the chip surface. We have to remove all covering layers using HF wet etching, after which we can easily recognise the rims of the diffusion regions that reveal the stored bit pattern. Some ROM technologies store bits not in the shape of the active area but by modifying transistor threshold voltages. In this case, additional dopant‑selective staining techniques have to be applied to make the bits visible. Together with an understanding of the (sometimes slightly scrambled) memory‑cell addressing, we obtain disassembler listings of the entire ROM content. Again, automated processing techniques can be used to extract the data from photos, but we also know cases where an enthusiastic smartcard hacker has reconstructed several kilobytes of ROM manually. Although ROM usually does not contain cryptographic keys, it often provides enough I/O and access‑control routines to design a non‑invasive attack – which is why reverse engineering remains the first step for most pirates.
Manual Microprobing: The most important tool for invasive attacks is a microprobing workstation. Its major component is a special optical microscope with a working distance of at least 8 mm between the chip surface and the objective lens. On a stable platform around a socket for the test package, we install several micropositioners, which allow us to move a probe arm with submicrometer precision over a chip surface. On this arm, we install a "cat whisker" probe – a metal shaft that holds a 10 µm diameter and 5 mm long tungsten‑hair, which has been sharpened at the end into a <0.1 µm tip. These elastic probe hairs allow us to establish electrical contact with on‑chip bus lines without damaging them. We connect them via an amplifier to a digital signal processor card that records or overrides processor signals and also provides the power, clock, reset, and I/O signals needed to operate the processor via the pins of the test package. A single well‑placed probe can unlock the entire memory bus if you know which lines carry the critical data. On the depackaged chip, the top‑layer aluminium interconnect lines are still covered by a passivation layer (usually silicon oxide or nitride), which protects the chip from the environment and ion migration. On top of this, we might also find a polyimide layer that was not entirely removed by HNO₃ but which can be dissolved with ethylenediamine. We have to remove the passivation layer before the probes can establish contact. The most convenient depassivation technique is the use of a laser cutter. The UV or green laser is mounted on the camera port of the microscope and fires laser pulses through the microscope onto rectangular areas of the chip with micrometre precision. Carefully dosed laser flashes remove patches of the passivation layer. The resulting hole in the passivation layer can be made so small that only a single bus line is exposed, preventing accidental contacts with neighbouring lines and the hole also stabilises the position of the probe and makes it less sensitive to vibrations and temperature changes. Complete microprobing workstations cost tens of thousands of dollars, with the more luxurious versions reaching over a hundred thousand US. The cost of a new laser cutter is roughly in the same region. Low‑budget attackers are likely to get a cheaper solution on the second‑hand market for semiconductor test equipment. With patience and skill it should not be too difficult to assemble all the required tools for even under ten thousand US by buying a second‑hand microscope and using self‑designed micropositioners. The laser is not essential for first results, because vibrations in the probing needle can also be used to break holes into the passivation.
Memory Read‑out Techniques: It is usually not practical to read the information stored on a security processor directly out of each single memory cell, except for ROM. The stored data has to be accessed via the memory bus where all data is available at a single location. Microprobing is used to observe the entire bus and record the values in memory as they are accessed. It is difficult to observe all (usually over 20) data and address bus lines at the same time. Various techniques can be used to get around this problem. For instance we can repeat the same transaction many times and use only two to four probes to observe various subsets of the bus lines. As long as the processor performs the same sequence of memory accesses each time, we can combine the recorded bus subset signals into a complete bus trace. Overlapping bus lines in the various recordings help us to synchronise them before they are combined. Every time you perform a firmware extraction via bus snooping, you rely on the processor repeating the same instruction stream predictably. In applications such as pay‑TV, attackers can easily replay some authentic protocol exchange with the card during a microprobing examination. These applications cannot implement strong replay protections in their protocols, because the transaction counters required to do this would cause an NVRAM write access per transaction. Some conditional‑access cards have to perform over a thousand protocol exchanges per hour and EEPROM technology allows only 10⁴‑10⁶ write cycles during the lifetime of a storage cell. An NVRAM transaction counter would damage the memory cells, and a RAM counter can be reset by the attacker easily by removing power. Newer memory technologies such as FERAM allow over 10⁹ write cycles, which should solve this problem. Just replaying transactions might not suffice to make the processor access all critical memory locations. For instance, some banking cards read critical keys from memory only after authenticating that they are indeed talking to an ATM. Pay‑TV card designers have started to implement many different encryption keys and variations of encryption algorithms in every card, and they switch between these every few weeks. The memory locations of algorithm and key variations are not accessed by the processor before these variations have been activated by a signed message from the broadcaster, so that passive monitoring of bus lines will not reveal these secrets to an attacker early. A poorly designed integrity check can become an open door for a skilled crack, as it forces the CPU to read every byte in a predictable order. Sometimes, hostile bus observers are lucky and encounter a card where the programmer believed that by calculating and verifying some memory checksum after every reset the tamper‑resistance could somehow be increased. This gives the attacker of course easy immediate access to all memory locations on the bus and simplifies completing the read‑out operation considerably. Surprisingly, such memory integrity checks were even suggested in the smartcard security literature, in order to defeat a proposed memory rewrite attack technique. This demonstrates the importance of training the designers of security processors and applications in performing a wide range of attacks before they start to design countermeasures. Otherwise, measures against one attack can far too easily backfire and simplify other approaches in unexpected ways. In order to read out all memory cells without the help of the card software, we have to abuse a CPU component as an address counter to access all memory cells for us. The program counter is already incremented automatically during every instruction cycle and used to read the next address, which makes it perfectly suited to serve us as an address sequence generator. We only have to prevent the processor from executing jump, call, or return instructions, which would disturb the program counter in its normal read sequence. Tiny modifications of the instruction decoder or program counter circuit, which can easily be performed by opening the right metal interconnect with a laser, often have the desired effect.
Particle Beam Techniques: Most currently available smartcard processors have feature sizes of 0.5‑1 µm and only two metal layers. These can be reverse‑engineered and observed with the manual and optical techniques described in the previous sections. For future card generations with more metal layers and features below the wavelength of visible light, more expensive tools additionally might have to be used. A focused ion beam (FIB) workstation consists of a vacuum chamber with a particle gun, comparable to a scanning electron microscope. Gallium ions are accelerated and focused from a liquid metal cathode with 30 kV into a beam of down to 5‑10 nm diameter, with beam currents ranging from 1 pA to 10 nA. FIBs can image samples from secondary particles similar to a SEM with down to 5 nm resolution. By increasing the beam current, chip material can be removed with the same resolution at a rate of around 0.25 µm³ nA⁻¹ s⁻¹. Better etch rates can be achieved by injecting a gas like iodine via a needle that is brought to within a few hundred micrometres from the beam target. Gas molecules settle down on the chip surface and react with removed material to form a volatile compound that can be pumped away and is not redeposited. Using this gas‑assisted etch technique, holes that are up to 12 times deeper than wide can be created at arbitrary angles to get access to deep metal layers without damaging nearby structures. By injecting a platinum‑based organo‑metallic gas that is broken down on the chip surface by the ion beam, platinum can be deposited to establish new contacts. With other gas chemistries, even insulators can be deposited to establish surface contacts to deep metal without contacting any covering layers. FIB editing is the ultimate tool for reverse engineering because it lets you rewrite the chip’s physical connectivity at will. Using laser interferometer stages, a FIB operator can navigate blindly on a chip surface with 0.15 µm precision, even if the chip has been planarised and has no recognisable surface structures. Chips can also be polished from the back side down to a thickness of just a few tens of micrometres. Using laser‑interferometer navigation or infrared laser imaging, it is then possible to locate individual transistors and contact them through the silicon substrate by FIB editing a suitable hole. This rear‑access technique has probably not yet been used by pirates so far, but the technique is about to become much more commonly available and therefore has to be taken into account by designers of new security chips. FIBs are used by attackers today primarily to simplify manual probing of deep metal and polysilicon lines. A hole is drilled to the signal line of interest, filled with platinum to bring the signal to the surface, where a several micrometre large probing pad or cross is created to allow easy access. Modern FIB workstations cost less than half a million US$ and are available in over hundred organisations. Processing time can be rented from numerous companies all over the world for a few hundred dollars per hour. Another useful particle beam tool are electron‑beam testers (EBT). These are SEMs with a voltage‑contrast function. Typical acceleration voltages and beam currents for the primary electrons are 2.5 kV and 5 nA. The number and energy of secondary electrons are an indication of the local electric field on the chip surface and signal lines can be observed with submicrometer resolution. The signal generated during e‑beam testing is essentially the low‑pass filtered product of the beam current multiplied with a function of the signal voltage, plus noise. EBTs can measure waveforms with a bandwidth of several gigahertz, but only with periodic signals where stroboscopic techniques and periodic averaging can be used. If we use real‑time voltage‑contrast mode, where the beam is continuously directed to a single spot and the blurred and noisy stream of secondary electrons is recorded, then the signal bandwidth is limited to a few megahertz. While such a bandwidth might just be sufficient for observing a single signal line in a 3.5 MHz smartcard, it is too low to observe an entire bus with a sample frequency of several megahertz for each line. Using an EBT, an attacker can dump entire bus transactions without physical contact, provided the clock is slowed down enough. EBTs are very convenient attack tools if the clock frequency of the observed processor can be reduced below 100 kHz to allow real‑time recording of all bus lines or if the processor can be forced to generate periodic signals by continuously repeating the same transaction during the measurement.
A processor is essentially a set of a few hundred flipflops (registers, latches, and SRAM cells) that define its current state, plus combinatorial logic that calculates from the current state the next state during every clock cycle. Many analog effects in such a system can be used in non‑invasive attacks. Some examples are: every transistor and interconnection have a capacitance and resistance that, together with factors such as the temperature and supply voltage, determine the signal propagation delays. Due to production process fluctuations, these values can vary significantly within a single chip and between chips of the same type. A flipflop samples its input during a short time interval and compares it with a threshold voltage derived from its power supply voltage. The time of this sampling interval is fixed relative to the clock edge, but can vary between individual flipflops. The flipflops can accept the correct new state only after the outputs of the combinatorial logic have stabilised on the prior state. During every change in a CMOS gate, both the p‑ and n‑transistors are open for a short time, creating a brief short circuit of the power supply lines. Without a change, the supply current remains extremely small. Power supply current is also needed to charge or discharge the load capacitances when an output changes. A normal flipflop consists of two inverters and two transmission gates (8 transistors). SRAM cells use only two inverters and two transistors to ground one of the outputs during a write operation. This saves some space but causes a significant short‑circuit during every change of a bit. A single clock glitch can crack open an authentication routine that otherwise would require months of cryptographic work. There are numerous other effects. During careful security reviews of processor designs it is often necessary to perform detailed analog simulations and tests and it is not sufficient to just study a digital abstraction. Smartcard processors are particularly vulnerable to non‑invasive attacks, because the attacker has full control over the power and clock supply lines. Larger security modules can be equipped with backup batteries, electromagnetic shielding, low‑pass filters, and autonomous clock signal generators to reduce many of the risks to which smartcard processors are particularly exposed.
Glitch Attacks: In a glitch attack, we deliberately generate a malfunction that causes one or more flipflops to adopt the wrong state. The aim is usually to replace a single critical machine instruction with an almost arbitrary other one. Glitches can also aim to corrupt data values as they are transferred between registers and memory. Of the many fault‑induction attack techniques on smartcards that have been discussed in the recent literature, it has been our experience that glitch attacks are the ones most useful in practical attacks. We are currently aware of three techniques for creating fairly reliable malfunctions that affect only a very small number of machine cycles in smartcard processors: clock signal transients, power supply transients, and external electrical field transients. Particularly interesting instructions that an attacker might want to replace with glitches are conditional jumps or the test instructions preceding them. They create a window of vulnerability in the processing stages of many security applications that often allows us to bypass sophisticated cryptographic barriers by simply preventing the execution of the code that detects that an authentication attempt was unsuccessful. Instruction glitches can also be used to extend the runtime of loops, for instance in serial port output routines to see more of the memory after the output buffer, or also to reduce the runtime of loops, for instance to transform an iterated cipher function into an easy to break single‑round variant. Power analysis is a form of reverse engineering that does not require opening the package – it reveals the instruction sequence from afar. Clock‑signal glitches are currently the simplest and most practical ones. They temporarily increase the clock frequency for one or more half cycles, such that some flipflops sample their input before the new state has reached them. Although many manufacturers claim to implement high‑frequency detectors in their clock‑signal processing logic, these circuits are often only simple‑minded filters that do not detect single too short half‑cycles. They can be circumvented by carefully selecting the duty cycles of the clock signal during the glitch. In some designs, a clock‑frequency sensor that is perfectly secure under normal operating voltage ignores clock glitches if they coincide with a carefully designed power fluctuation. We have identified clock and power waveform combinations for some widely used processors that reliably increment the program counter by one without altering any other processor state. An arbitrary subsequence of the instructions found in the card can be executed by the attacker this way, which leaves very little opportunity for the program designer to implement effective countermeasures in software alone. Power fluctuations can shift the threshold voltages of gate inputs and anti‑tampering sensors relative to the unchanged potential of connected capacitances, especially if this occurs close to the sampling time of the flipflops. Smartcard chips do not provide much space for large buffer capacitors, and voltage threshold sensors often do not react to very fast transients. In a potential alternative glitch technique that we have yet to explore fully, we place two metal needles on the card surface, only a few hundred micrometres away from the processor. We then apply spikes of a few hundred volts for less than a microsecond on these needles to generate electrical fields in the silicon substrate of sufficient strength to temporarily shift the threshold voltages of nearby transistors.
Current Analysis: Using a 10‑15 Ω resistor in the power supply, we can measure with an analog/digital converter the fluctuations in the current consumed by the card. Preferably, the recording should be made with at least 12‑bit resolution and the sampling frequency should be an integer multiple of the card clock frequency. Drivers on the address and data bus often consist of up to a dozen parallel inverters per bit, each driving a large capacitive load. They cause a significant power‑supply short circuit during any transition. Changing a single bus line from 0 to 1 or vice versa can contribute in the order of 0.5‑1 mA to the total current at the right time after the clock edge, such that a 12‑bit ADC is sufficient to estimate the number of bus bits that change at a time. SRAM write operations often generate the strongest signals. By averaging the current measurements of many repeated identical transactions, we can even identify smaller signals that are not transmitted over the bus. Signals such as carry bit states are of special interest, because many cryptographic key scheduling algorithms use shift operations that single out individual key bits in the carry flag. Even if the status bit changes cannot be measured directly, they often cause changes in the instruction sequencer or microcode execution, which then cause a clear change in the power consumption. The various instructions cause different levels of activity in the instruction decoder and arithmetic units and can often be quite clearly distinguished, such that parts of algorithms can be reconstructed. Various units of the processor have their switching transients at different times relative to the clock edges and can be separated in high‑frequency measurements.
Insert random‑time delays at the clock‑cycle level using a hardware random bit generator and a pseudo‑random generator. This makes it difficult for attackers to predict instruction execution times. To prevent reconstruction of the internal clock from current consumption, the processor should exhibit activity even during inactive periods (e.g., by performing dummy writes).
Design a multithreaded architecture that schedules threads randomly at the instruction level. Multiple register sets make the execution flow non‑deterministic, hindering reverse engineering.
A sensor that triggers if no clock edge is seen for a defined time limit (e.g., 0.5 µs) and immediately grounds all buses and registers. The sensor should be integrated into the normal reset sequence, making it difficult to disable via FIB editing.
Permanently destroy test logic by placing critical test pads in the scribe lines cut during wafer dicing. This prevents attackers from using test circuits to dump memory.
Replace a full 16‑bit PC with a segment register + 7‑bit offset counter. The offset resets after 127 bytes, forcing jumps in code and preventing sequential read‑out of entire memory via bus snooping.
Add metallisation meshes above the circuit that are continuously monitored for interruptions. While not foolproof, they make manual microprobing more difficult and force attackers to use FIB drilling, increasing the cost of attacks.
No single countermeasure is perfect; a combination of techniques is required to raise the bar against mass‑market pirates. However, fully invasive attacks with FIB tools remain difficult to stop in the smartcard form factor due to the lack of battery‑backed zeroisation.
We have presented a basis for understanding the mechanisms that make microcontrollers particularly easy to penetrate. With the restricted program counter, the randomised clock signal, and the tamper‑resistant low‑frequency sensor, we have shown some selected examples of low‑cost countermeasures that we consider to be quite effective against a range of attacks. There are of course numerous other more obvious countermeasures against some of the commonly used attack techniques which we cannot cover in detail in this overview. Examples are current regulators and noisy loads against current analysis attacks and loosely coupled PLLs and edge barriers against clock glitch attacks. A combination of these together with e‑field sensors and randomised clocks or perhaps even multithreading hardware in new processor designs will hopefully make high‑speed non‑invasive attacks considerably less likely to succeed. Other countermeasures in fielded processors such as light and depassivation sensors have turned out to be of little use as they can be easily bypassed. No single countermeasure will stop a nation‑state‑level reverse engineering effort, but a layered approach can deter mass‑market pirates. We currently see no really effective short‑term protection against carefully planned invasive tampering involving focused ion‑beam tools. Zeroisation mechanisms for erasing secrets when tampering is detected require a continuous power supply that the credit‑card form factor does not allow. The attacker can thus safely disable the zeroisation mechanism before powering up the processor. Zeroisation remains a highly effective tampering protection for larger security modules that can afford to store secrets in battery‑backed SRAM (e.g., DS1954 or IBM 4758), but this is not yet feasible for the smartcard package. Ultimately, the battle between defenders and hackers will continue, as each new defensive layer inspires a new wave of firmware extraction and unlock techniques.