Months ago, in the article When Logic Meets Physics, I wrote that I was taking part in a competition whose goal was to develop a RISC-V-based processor. At that time, we were much closer to the fundamental building blocks and the design flow than to a complete CPU: writing synthesizable hardware, validating circuits, and following the path between a logical description and something that might eventually become silicon. A few months later, here we are: the Design Microchip Supreme — DMS team. The Supreme was Pedro’s idea. Sorry, Pedro, but I still think Supreme ruined the name.

At the time of writing, we have not finished developing this phase’s chip and, for fairly obvious reasons, I do not intend to turn this article into public documentation of our project while the competition is still underway. The goal here is different: to record some of the decisions, difficulties, discoveries, and curiosities that appeared when concepts I already knew fairly well had to stop existing as relatively isolated pieces and start working together inside the same microarchitecture. Before the competition, I already knew about ALUs, registers, multiplexers, state machines, buses, ISAs, Assembly, and many of the problems that appear when software meets hardware; working with microcontroller firmware practically forces you to live with much of that. What changed was not discovering that these pieces exist. It was having to make them agree with one another, cycle after cycle, until the whole thing actually behaved like a processor.

Before ADD was ADD

There is something curious about the history of computing: the desire to build machines capable of performing mathematical operations is much older than the digital processor as we know it today. For centuries, mechanical calculators based on gears, wheels, and levers appeared; later, analog computers represented mathematical quantities through voltages, currents, and other physical phenomena. This is where the name operational amplifier becomes especially interesting. These circuits were widely used in analog computers to perform operations such as addition, integration, and differentiation. Before ADD was a sequence of bits entering a decoder, an addition could literally emerge from the electrical behavior of a circuit configured to perform it.

I like this history because it takes away some of the magic we place around modern computers. At the core, we are still trying to make matter represent and manipulate mathematical abstractions; what changed dramatically was the level of abstraction available to the people designing these machines. Today I can write something as innocent as result = a + b; in Verilog and let a huge toolchain transform that description into an implementation made of much simpler digital elements. The addition does not stop existing physically. I simply do not have to draw, by hand, every gate that will make it exist.

I thought I would have to draw everything by hand

This was one of my first surprises in the competition. I already knew what an HDL was, knew Verilog, and understood the role of synthesis, but I still carried an exaggerated image of processor development as spending an absurd amount of time building the circuit almost gate by gate: multiplexers, flip-flops, adders, comparators, and wires until something resembling a CPU emerged. I expected suffering. I received suffering. I was simply wrong about its type.

HDLs such as Verilog exist precisely to let us work at a much higher level of abstraction. Instead of individually drawing every gate required for an operation, we describe behavior, signal relationships, registers, combinational logic, and state machines, and synthesis tools turn that description into a hardware network. This makes life considerably easier, but it brings a curious trap: Verilog looks enough like a programming language to occasionally try to convince you that it is software. Then the hardware reminds you. An if does not necessarily mean a decision that will execute after the previous line, as it would in C; two blocks can represent circuits that exist and react simultaneously; a variable can represent a connection or state; and code that is perfectly valid syntactically can still synthesize into a circuit quite different from the one you imagined.

The change for me was less about learning new components and more about continuously thinking in terms of structure, simultaneity, and state. The code is not simply giving orders to a finished machine. Much of the time, it is describing the machine itself.

So, what is a CPU?

If I had to define a Central Processing Unit, or CPU, in deliberately simple terms, I would say that it is a programmable machine capable of fetching an instruction, interpreting what that instruction means, executing the corresponding operation, and repeating the process. Conceptually, that does not sound too frightening. So let us imagine an extremely sophisticated processor called the Yuri Processing Unit, or YPU. For now, the YPU understands only two instructions: LOAD and ADD.

LOAD VALUE_1
LOAD VALUE_2
ADD VALUE_1, VALUE_2

The expected behavior seems obvious: load VALUE_1, load VALUE_2, and add them. Done. We have a processor. That explanation does not last long once we start asking inconvenient questions. Where is VALUE_1? When the CPU reads a value, where does it store it? How does it know that a particular sequence of bits means ADD? How does it decide which operands enter the operation? Where does it put the result? How does it know where the next instruction is? And who coordinates all of this so that no part decides to do the right thing at the wrong time?

flowchart LR
    PC["Program Counter"] --> IMEM["Instruction memory"]
    IMEM --> DEC["Decoder / Control"]
    IMEM --> RF["Registers"]
    RF --> MUX["Operand selection"]
    DEC --> MUX
    DEC --> ALU["ALU"]
    MUX --> ALU
    ALU --> RF
    ALU --> PC

A deliberately simplified view. The goal is not to represent the team’s actual microarchitecture, but to show how the boxes begin to depend on one another.

This is where the pretty boxes in architecture diagrams stop being just boxes. We need registers, data paths, a unit capable of performing operations, some way to access memory, a counter to track execution flow, logic to interpret the bits of each instruction, and control signals telling each block what should happen at that moment. “Load two numbers and add them” remains a simple idea. Making a machine repeat it correctly for dozens of different instructions without destroying its own state along the way is another story.

Multiplexers: choosing is also computing

One component that appears constantly along these paths is the multiplexer, or MUX. The idea is simple: several possible inputs exist, and a selection signal decides which one is forwarded to the output. If the ALU can receive either a register value or an immediate extracted from the instruction as its second operand, for example, some circuit has to decide which of those reaches it in that cycle; that circuit may be a multiplexer. The same reasoning appears when choosing the next Program Counter value, the source of data written to the register file, or which result should continue through the datapath.

Multiplexing, however, is far from being an idea exclusive to processors. Imagine a system with four analog sensors and only one analog-to-digital converter: instead of keeping four ADCs, we can select which sensor is connected to the converter at each moment. An audio or video input selector follows a similar idea. In telecommunications, different forms of multiplexing allow multiple streams to share the same physical medium. The implementation changes, the scale changes, and sometimes even the domain changes from digital to analog, but the fundamental question remains similar: among the several things that could pass through here, which one should pass now?

flowchart TB
    subgraph Real_world["Outside the CPU"]
        S1["Sensor A"] --> AMUX{"Analog MUX"}
        S2["Sensor B"] --> AMUX
        S3["Sensor C"] --> AMUX
        SEL1["Selection"] --> AMUX
        AMUX --> ADC["ADC"]
    end

    subgraph CPU["Inside the CPU"]
        REG["Register value"] --> DMUX{"MUX"}
        IMM["Immediate"] --> DMUX
        PCV["Program Counter"] --> DMUX
        CTRL["Control"] -->|selection| DMUX
        DMUX --> ALU["ALU"]
    end

That may be why I started seeing MUXes as a kind of scorecard for the architecture. In the diagram, they look like small selectors between paths. In practice, each one exists because at some point we had to make a decision about data flow. And every new possible path is also a new opportunity for the selection signal to be wrong.

Registers: short-term memory, state, and time

If multiplexers choose paths, registers allow the CPU to preserve state. In simple terms, a register is a collection of elements capable of storing bits. Some are defined by the ISA and visible to programmers, such as general-purpose registers; others have special functions, such as the Program Counter; and some exist only because of a microarchitectural decision, holding intermediate results, control state, or data that needs to survive until the next cycle. What they have in common is that they allow the circuit to remember something after the surrounding combinational logic has changed.

Talking about state forces us to talk about the clock. In a synchronous circuit, we generally want the main state changes to happen in coordination with a clock edge. Combinational logic can spend part of the cycle calculating the next value, but the register captures that value only when the selected edge occurs, often the rising edge. In many register files, for example, writes may be synchronous while reads are combinational, depending on the implementation. There are also asynchronous signals, such as certain resets, that can affect state independently of a clock edge. This distinction matters because a synchronous system gains predictability by organizing most changes at well-defined instants; anything arriving asynchronously must be handled carefully to avoid hard-to-reproduce behavior or timing problems.

In Verilog, this appears quite concretely. In clock-triggered sequential logic, non-blocking assignments using <= are usually the right choice because they express state updates that should occur in coordination after that event is evaluated. In combinational logic, blocking assignments using = generally express the calculations that form the block’s combinational value more clearly. This is not just style: confusing these models can make the simulation represent an order that does not match the hardware intent.

always @(posedge clk) begin
    if (reset)
        pc <= 32'b0;
    else
        pc <= next_pc;
end

This does not mean “execute this line after the previous one” as in a traditional program. It describes state: on the rising edge of the clock, the register associated with pc captures a new value. That is a small syntactic difference and a huge difference in how you think. When several registers do this together, the entire circuit seems to take a small step through time.

An instruction is just a number with self-esteem

In RISC-V, an RV32 instruction has 32 bits. To us, it may appear as add x5, x6, x7; to the processor, it is simply a bit vector whose meaning depends on how each field is interpreted. An R-type instruction, for example, contains fields such as funct7, rs2, rs1, funct3, rd, and opcode. The opcode helps identify the general instruction class; rs1 and rs2 identify source registers; rd points to the destination; and funct3 and funct7 help distinguish operations that share part of the encoding.

31..2524..2019..1514..1211..76..0
funct7rs2rs1funct3rdopcode

When you look at this only as an ISA, it is an instruction format. When you have to implement the decoder, every field starts to mean bit extraction, comparisons, and control signals that must agree with the rest of the datapath. The interesting part was not discovering that these fields exist — I already knew them — but realizing how directly certain ISA choices influence the most convenient way to build the circuit that interprets them.

The register that wins by doing nothing

One of my favorite small RISC-V decisions is x0. It is a register whose value is permanently zero. You can try to write to it as much as you like; it will remain zero. addi x0, x0, 42 is a particularly sophisticated way of failing to store 42 anywhere. Reserving a register to store nothing may initially seem wasteful, but having zero always available simplifies several operations, allows results to be discarded simply by using x0 as the destination, and helps represent several pseudo-instructions without unnecessarily expanding the ISA.

For someone writing Assembly, this becomes a convenience. For someone implementing the register file, it becomes a concrete rule: reads from x0 must produce zero, and writes targeting x0 must be ignored. It is a good example of something that appears repeatedly in this kind of project: a short sentence in a specification can become a very specific responsibility inside the hardware.

funct7, or how seven bits also gain personality

Another detail that gains weight during implementation is funct7. Some operations share the same opcode and part of the same fields, so the decoder has to inspect combinations of opcode, funct3, and funct7 to discover what is actually being requested. Decoding an instruction is therefore far more than turning a number into a pretty operation name. The decision must become concrete actions: which registers will be read, which source will feed the ALU, which operation the ALU should perform, whether the register file will be written, whether memory will be accessed, and how the Program Counter will be updated.

That is where a small mistake acquires large consequences. Misinterpreting a few bits does not merely produce a wrong label; it can make the right data take the wrong path, or the wrong data take the right path, which is a particularly polite kind of bug because it sometimes produces a plausible result before ruining your afternoon.

Who scattered the immediate bits on the floor?

Then there are immediates. Some instructions carry part of their operands inside their own 32 bits, and it would be intuitive to imagine those bits always sitting together, in order, waiting to be read. In some formats the situation is friendly. In others, especially the formats used by branches and jumps, the first impression is that someone dropped the immediate on the floor and put the pieces back wherever they found room.

In a branch, for example, the immediate is reconstructed from fields distributed across the instruction. This organization looks strange when viewed by a human looking at the encoding, but it has an important architectural reason: keeping fields such as rs1, rs2, and funct3 in consistent positions simplifies other decoding paths, while certain bits are positioned to support the implementation and encoding of offsets. What looks disorganized from one perspective can be a deliberate choice from another.

3130..2524..2019..1514..1211..876..0
imm[12]imm[10:5]rs2rs1funct3imm[4:1]imm[11]opcode

Yes, imm[11] really is there.

That was a recurring feeling throughout the competition: concepts I already knew gained a second reading when they had to enter a real circuit. Before, I looked at an instruction format and saw the specification. Now, at the same time, I also see sign extension, concatenation, comparators, MUXes, and control signals that must produce exactly that behavior.

The same 11111111 can have two personalities

Another source of amusement is the difference between signed and unsigned values. Physically, a register does not store “a negative number” using a different material from “a positive number”; it stores bits. Meaning comes from interpretation. In 8 bits, 11111111 can represent 255 when treated as unsigned or -1 when interpreted as signed two’s complement. The vector is the same. The semantics change, and that change begins spreading through the circuit when the ISA contains operations that distinguish the two cases.

Signed and unsigned comparisons require different reasoning. Extensions have to preserve the right meaning. An arithmetic shift is not the same as a logical shift. Multiplication makes this distinction particularly visible. Multiplying two 32-bit operands can require a full 64-bit product; when we need only the lower half, some signed and unsigned cases produce the same lower 32 bits. When we want the upper half, however, operand interpretation matters directly. Instructions such as MULH, MULHU, and MULHSU exist for different signed and unsigned combinations, and a detail that initially seems semantic suddenly becomes a concrete decision about sign extension, operand width, and which result bits are preserved.

This was one of those problems that first had to be correct in the logic of my head before it could be correct in Verilog. Writing $signed or $unsigned does not fix an unclear intention. First I need to know the width of each operand, how it should be extended, and which interpretation I want to preserve. Verilog also offers additional opportunities to stumble: the signedness of expressions can change as concatenations, slices, constants, and casts enter the picture. Relying too heavily on implicit conversions is an efficient way to write something that passes obvious tests and fails precisely on the boundary value nobody remembered to try. In this part of the project, making width and signedness explicit stopped feeling verbose and started acting as documentation for the circuit I actually wanted to build.

Four multiplications and one register

It was precisely in the multiplication instructions that one of the architectural decisions I liked most appeared. A quick note: multiplication is not part of pure RV32I; it belongs to RISC-V’s M extension. During implementation, I could take a more expensive hardware path and keep independent storage structures associated with the different multiplication variants, or share a single fixed-size register to hold the full result used by those operations. In very simple terms, I could spend more resources and keep each path more independent, or save hardware and increase the responsibility of the control logic.

I chose to share the register. If the different operations do not need to keep their results simultaneously, replicating storage can mean spending area without an immediate need. But “use less hardware” never means “get something for free.” A shared resource has to be arbitrated and coordinated. The FSM must guarantee that nobody consumes an old value or overwrites the resource at the wrong time; the decision may limit future parallelism; and an optimization that looks excellent when considering only the number of registers may become less attractive once timing, routing, power, or the evolution of the microarchitecture enter the discussion.

This decision summarized a change in perspective that the project reinforced for me. The question is rarely just “which solution works?” It is usually “which trade-off do we want to make?” Area, power, frequency, control complexity, ease of verification, and room for expansion compete with one another. A finished diagram shows the result of the choice; building the CPU makes you live through the reasons that choice had to exist.

Datapath: where the bits travel

So far I have talked about ALUs, registers, multiplexers, immediates, and several other elements as recognizable pieces. When we start looking at the paths that actually carry, transform, and store data between those pieces, we reach the datapath. In simple terms, it is the part of the microarchitecture through which values move: operands leave the register file, pass through selectors, enter the ALU, may access memory, return to a register, and, depending on the instruction, even influence the next Program Counter value. If the ISA describes what an instruction means, the datapath helps answer where the data must travel for that meaning to happen.

An analogy that works reasonably well is a city with roads already built. Data is the traffic; registers, memory, and the ALU are places it can pass through; multiplexers are intersections that choose routes. But the roads do not know where each car needs to go. Something still has to say which paths are open, when a value may be stored, and which operation should happen at that moment.

The FSM is the conductor nobody sees

In our implementation, much of this coordination is handled by a finite state machine, or FSM. An FSM represents control as a finite set of states and transition rules. In a multicycle CPU, this is especially useful because an instruction may require different actions over several cycles: at one point we fetch or prepare an instruction, at another we interpret its fields, then we configure the paths needed in the datapath, and at some point we register the result. I will not expose our implementation’s actual state sequence here, but the important idea is that some logic knows which point in execution we are at and, based on that, which actions may happen.

flowchart LR
    FSM["FSM / Control unit"] -->|"selectors, enables, operation"| DP["Datapath"]
    DP -->|"flags, comparisons, completion"| FSM

    subgraph D["Datapath"]
        PC2["PC"] --> MUX2{"MUXes"}
        RF2["Registers"] --> MUX2
        MUX2 --> ALU2["ALU"]
        ALU2 --> RF2
    end

The FSM and datapath work as two halves of the same conversation. The datapath contains the resources and paths capable of performing operations; the FSM produces control signals that select those paths, enable writes, choose operations, and determine when state may advance. The datapath, in turn, returns information that can influence control, such as comparison results, branch conditions, or an indication that an operation is complete. When the two agree, an instruction progresses in an orderly way. When they do not, you open the waveform.

Here is a conceptual view of how a multicycle FSM might organize execution. It does not represent the actual states in our implementation; it simply shows that the CPU does not have to do everything at once.

stateDiagram-v2
    [*] --> Fetch
    Fetch --> Decode
    Decode --> Execute
    Execute --> Finish
    Finish --> Fetch

“It works on my side”

There is a particularly dangerous sentence in any project developed by several people: “it works on my side.” There are four of us building parts of a system that eventually has to behave as one machine. That means it is not enough for each module to be correct in isolation. Two people can implement perfectly valid modules and still produce an incorrect system together. An interface may be interpreted in two different ways; someone may assume a signal remains valid for an entire cycle while another part expects only a pulse; an apparently local change may break an old assumption in another block. Sometimes the bug is not on either side of the connection. It is exactly in the way the two sides understand the connection.

Perhaps because of my software background, one of the things I invested in early was the infrastructure around the processor. I set up a reproducible virtual environment with the tools needed to develop and test the project, reducing the classic “it works on my machine” as much as possible. We organized the project in Git and created a CI pipeline that automatically runs a testbench suite. Every relevant change retests functionality that already existed. If a new instruction silently breaks something that worked two weeks ago, we want to discover it in the commit that introduced the problem, not after fifteen more changes have been stacked on top of it.

flowchart LR
    DEV["Project change"] --> GIT["Git"]
    GIT --> CI["CI"]
    CI --> TB["Testbench suite"]
    CI --> FW["Assembly test firmware"]
    TB --> CHECK{"Passed?"}
    FW --> CHECK
    CHECK -->|Yes| OK["Keep going"]
    CHECK -->|No| WAVE["Open the waveform and find the liar"]

If it can break, we try to break it

In addition to specific testbenches, we also maintain an Assembly file that acts as validation firmware for the processor itself. It grows together with the CPU. Every new instruction or implemented feature also enters this program, and the goal is not merely to test the happy path. We try to combine behaviors, reuse results, force boundary values, and create sequences that may expose unexpected interactions. An addition can work perfectly in isolation and fail after a certain sequence; a branch can pass an obvious case and break when it depends on a previous result; a decoder change can fix one instruction family and introduce a regression in another.

Over time, the question stops being only “does this instruction work?” and becomes “how many different ways can we make it fail?” When building something so dependent on state, timing, and interaction between blocks, testing each module in isolation is only the first layer. The interesting behavior appears when instructions start living together. The regression suite and test firmware therefore became part of development, rather than a separate stage that happens after somebody declares the processor “ready.”

This was also a fun part of the boundary between software and hardware. Git, automation, continuous integration, and regression tests do not belong exclusively to the software world; the object being tested is different, but the principle remains extremely valuable. Every change should bring evidence that the expected behavior still exists.

Where the truth hides among dozens of signals

When something inevitably breaks, one of the most characteristic activities begins: opening the waveforms and trying to reconstruct, cycle by cycle, what the processor believed it was doing. Was the Program Counter pointing to the correct instruction? Did memory provide the expected bits? Did the decoder interpret opcode, funct3, and funct7 correctly? Did rs1 and rs2 point to the right registers? Was the immediate reconstructed correctly? Was the FSM in the expected state? Did the ALU receive the correct operands? Did the result return to rd? And, perhaps more importantly than all of those questions, did it happen in the cycle when it was supposed to happen?

That last point may look trivial when written this way, but it is one of the most important differences between thinking only about an operation’s logic and thinking about hardware. It is not enough for a value to be correct; it must be correct at the right time. A perfect signal arriving one cycle late is still a bug. A correct result written too early is also a bug. A state change at the wrong instant can produce a consequence that appears several instructions later. Waveforms have an almost investigative quality: you begin with the crime — the firmware produced an impossible result — and work backward looking for the first instant when the processor’s reality diverged from the reality you expected. At some point, one bit made a questionable decision. The job is to find which one.

Knowing the pieces is not building the system

Before the competition, I already knew what an ALU, register file, Program Counter, control unit, multiplexers, and state machines were. I had studied instruction formats, written Assembly, worked directly with microcontroller registers, and spent enough time with hardware to understand many of these concepts. The value of that experience was not introducing me to a new list of components; it was forcing me to see the relationships between components I already knew. An isolated ALU is relatively easy to understand. An isolated decoder is too. A register file, MUX, or FSM is similar. The difficulty appears when the correct instruction has to be fetched, interpreted, sent to the right operands, executed by the right unit, have its result delivered to the right destination, and update the machine’s state at the precise moment — so that everything can immediately start again without carrying an error from the previous cycle.

Perhaps this is the idea that best summarizes what I have learned so far: a processor is not difficult merely because it has many pieces; it is difficult because all those pieces have to agree about reality at the same time. The decoder has to interpret the instruction the same way the datapath expects; the register file has to provide the data the selected paths actually consume; the FSM has to know when each write is allowed; the Program Counter has to advance or branch according to what just happened; and the verification environment has to notice when any of those relationships stops being true. I already knew the instruments. The competition put me in the middle of the orchestra.

Once the RTL works, physics starts charging interest

The CPU is not finished yet, and making the logical behavior work is only part of the journey. One of the next steps is to turn the knowledge currently concentrated in RTL, comments, testbenches, waveforms, and the heads of the people who made the decisions into real documentation: interfaces, signals, relevant states, data flow, conventions, architectural decisions, and the reasons behind them. Documentation is not as exciting as adding a new instruction, but a system that can only be understood by the people who were present while it emerged is already accumulating dangerous debt.

After that comes the part that directly connects this text to the previous article: taking the project through synthesis and physical implementation in OpenLane. So far, much of our reasoning has focused on functional correctness: does the right instruction produce the right result? Does the FSM advance correctly? Do the testbenches pass? Once the project enters the physical flow, those questions remain valid but are no longer enough. We start worrying about floorplanning, placement, clock-tree synthesis, routing, parasitic extraction, and timing analysis. The circuit must not only be correct; it must fit, be routable, and operate within the frequency and power constraints we choose.

flowchart LR
    RTL["RTL / Verilog"] --> SYN["Synthesis"]
    SYN --> FP["Floorplan"]
    FP --> PLC["Placement"]
    PLC --> CTS["Clock tree"]
    CTS --> RT["Routing"]
    RT --> RC["RC extraction"]
    RC --> STA["Timing / STA"]
    STA --> PPA["Area, power, and performance"]

This is where physics returns many of the abstractions that Verilog kindly hid from the project. Interconnects have resistance and capacitance; cells have delays; fanout and load matter; the clock does not magically arrive everywhere at the same instant; process, voltage, and temperature variations change behavior; and parasitics that did not appear in the RTL description enter the calculation after layout. We also have to consider power-related effects such as voltage drops and the impact that less-than-ideal electrical conditions can have on margins and performance. Two blocks that looked like perfect neighbors in the code may end up physically separated by a distance with real consequences.

One of the most direct consequences is timing. Between two clock edges there is a finite amount of time for a value to travel through combinational logic and become stable before being captured by the next sequential element. If the path is too slow, setup violations may occur; if the timing relationship after the edge is not respected, we may have hold violations. This creates a situation I particularly like as a summary of digital design: a circuit can be logically correct and physically wrong. The testbench passes, the Boolean function is right, and yet the implementation may not operate reliably at the frequency you wanted.

We will also have to look much more carefully at performance, area, power, and energy efficiency. Increasing frequency can make critical paths harder and raise consumption. Parallelizing hardware can reduce the number of cycles while increasing area, capacitance, and switching activity. Sharing resources — as in the register used by the multiplication instructions — can save silicon, but it can also create more complex paths or reduce opportunities for parallelism. A choice that looks clearly efficient when viewed only at the RTL may receive a second interpretation once area, timing, and power reports enter the conversation. That is precisely where today’s decisions stop being merely elegant ideas and begin receiving a grade from physics.

That is why I still consider this text a report from the middle of the road. Some decisions that look good today will probably be questioned once synthesis, timing, area, and power reports are in front of us. A critical path may reveal that we were too optimistic. A structure may occupy more area than expected. An apparently irrelevant routing detail may force an architectural revision. Some testbench will still find a combination nobody anticipated, some waveform will still make me reconsider my life choices, and there is a statistically significant chance that Pedro will try to put Supreme in the name of something else.

When I entered the competition, I was not starting from zero in computer architecture. I already knew many of the pieces and had lived with many of the concepts in this article. The experience changed something else: how I see the whole system. Verilog abstracts away an enormous amount of the work required to build a circuit gate by gate, but it does not abstract away engineering decisions. In fact, when ALUs, registers, MUXes, FSMs, decoders, timing, and verification all have to work together, those decisions become more visible than ever.

There is also a next step that makes all of this especially concrete: in the next phase of the competition, we will take the project to a real FPGA. So far, the processor has lived mainly in RTL, testbenches, and waveforms. On the FPGA, we will synthesize and implement the circuit in reconfigurable hardware, connect the required signals, and watch our processor actually execute. This is the moment when what is currently a hardware description begins to gain a physical life.

FPGA stands for Field-Programmable Gate Array: a chip full of logic blocks, registers, memory, and interconnects that can be configured after manufacturing. Instead of merely simulating a CPU’s behavior on a computer, we load a configuration that physically turns those resources into a CPU running in parallel, with a clock, electrical signals, and real limitations. An FPGA matters because it occupies the space between an RTL description and a manufactured chip: it lets us test the architecture in hardware, observe the complete system, and change the design without fabricating new silicon.

There is something almost absurd about it: the same piece of matter can become a processor, a video controller, a communication interface, or a specialized accelerator depending on the configuration we load. We are not merely running a program on finished hardware; we are using a program to define what hardware will exist. It is a slightly crazy form of programming, because the result is not just a sequence of instructions — it is an entire machine taking shape.

And it can get much more complicated

Even after making a multicycle CPU work, we are still far from exhausting the possibilities of an architecture. A next step could be dividing execution into a pipeline, allowing different instructions to occupy different stages at the same time. This can increase throughput, but it also creates new problems: dependencies between instructions, data hazards, control hazards caused by branches, and the need for mechanisms such as stalls, forwarding, or branch prediction.

It would also be possible to add multiple cores, which would introduce questions about communication, synchronization, shared memory, and cache coherence. Speaking of caches, a memory hierarchy could reduce the cost of frequent accesses, but it would require handling misses, replacement policies, writes, and data consistency. From there, an MMU could translate virtual addresses into physical addresses using page tables and a TLB to keep recent translations nearby. Virtual memory, process protection, and different privilege levels would add even more responsibility to the hardware.

The processor also needs to communicate with the outside world. Adding I/O interfaces, controllers, interrupts, and GPIO would let the CPU interact with real peripherals and devices, but it would require defining protocols, addressing, timing, and safe ways to cross the boundary between hardware and software. Every one of these extensions is entirely possible; none of them is merely “one more block.” They add states, paths, rules, and interactions that must remain coherent.

It is easy to look at a simple CPU and imagine that we only need to keep stacking features until we reach a modern processor. In practice, every layer makes the previous ones harder to verify, optimize, and explain. A pipeline affects hazards; hazards affect control; caches affect memory behavior; the MMU affects addresses; I/O and interrupts affect execution flow. The architecture can become much more complex — and that escalation is precisely what makes it so interesting to realize how much a few pieces are already capable of teaching us.

In the end, 32 bits seem like very little.

Until you have to decide what each one means — and then convince physics to agree.