Linux Kernel PCI Data Structures-Free Linux Device Drivers Training Online

PREV_LEC  |  NEXT_LEC

Linux Kernel PCI Data Structures
Free Linux Kernel Development Course — PCI Subsystem Internals, struct pci_dev, and Modern IRQ Handling on Kernel 6.x
Part 5 of PCI Series
Kernel 6.x APIs
Hands-on Driver Demo

If you are following our free Linux kernel development course, you already know how a PCI device gets discovered on the bus. In this lecture we go one level deeper into the Linux kernel PCI data structures that every PCI/PCIe driver author must understand before writing a single line of probe() code. We will walk through the three software layers that bring a PCI device to life inside the kernel, then dissect struct pci_dev field by field using the current mainline kernel source, and finish by writing an original driver that dumps this information live from dmesg. This is exactly the kind of practical, from-scratch content you’d expect from a free linux device drivers course — no copy-pasted book listings, only concepts explained in plain language and code you can build yourself.

PCI subsystem struct pci_dev pci_alloc_irq_vectors MSI / MSI-X free embedded linux course

What You Will Learn

  • The three software layers of the Linux PCI stack
  • Every important field of struct pci_dev on kernel 6.x
  • How the irq field changes meaning under legacy INTx vs MSI/MSI-X
  • Writing an original driver that reads and prints pci_dev fields
  • Common mistakes when accessing PCI device state

Prerequisites

This lecture continues our PCI/PCIe device driver series. Before reading further, we recommend finishing the earlier lectures on PCI/PCIe bus topology and enumeration, header types and BDF addressing, PCI address spaces and BAR mapping, and PCI interrupt distribution (legacy INTx vs MSI/MSI-X). A working knowledge of C, basic kernel module loading (insmod/rmmod), and a QEMU test environment is assumed — all covered earlier in this free linux kernel development course.

The Three Layers of the Linux PCI Software Stack

One of the most confusing parts for newcomers to Linux kernel PCI data structures is understanding *who* actually creates the device objects the driver eventually binds to. The kernel splits this responsibility across three cooperating layers, each with a distinct job:

Linux PCI Software Stack
Layer 1: Architecture PCI Glue  →  kicks off bus scan on boot
Layer 2: Host Controller (root complex)  →  SoC/board specific, config-cycle callbacks
Layer 3: PCI Core (drivers/pci/)  →  builds bus/device tree, creates struct pci_dev
Layer 4: Endpoint Driver (your code)  →  struct pci_driver, probe()/remove()

Architecture PCI glue is the earliest piece to run. On modern ARM64 and most embedded SoCs this work has largely moved into the generic PCI core plus the host controller driver, but conceptually its job stays the same: get the bus numbered and ready before anything else touches it.

The host controller (often called the root complex) is SoC-specific. You will find these drivers under drivers/pci/controller/ in a modern kernel tree — for example the generic DesignWare-based controllers live under drivers/pci/controller/dwc/, while other SoC vendors ship their own controller driver in the same directory. This layer is responsible for generating configuration-space read/write cycles, exposing the memory/IO resource windows the SoC provides, and wiring up the INTx and MSI interrupt lines. Without a correctly initialized host controller, the PCI core has no bus to scan at all.

The PCI core, implemented mainly in drivers/pci/probe.c, is where the actual struct pci_dev and struct pci_bus objects are born. The core walks every bus/device/function combination reachable through the host controller, reads the standard configuration header from each function that responds, and allocates a pci_dev instance for it. It also creates the sysfs entries under /sys/bus/pci/devices/ that you use every day with tools like lspci, handles PCI Express port/bridge devices, and provides the MSI/MSI-X interrupt allocation framework your driver will call into.

Once the core has built this tree, it matches each pci_dev against the pci_device_id tables registered by loaded pci_driver structures and calls your probe() function — which brings us to the object your driver actually works with: struct pci_dev.

Inside struct pci_dev on a Modern Kernel

Every PCI or PCIe function detected on the bus gets exactly one struct pci_dev, defined in include/linux/pci.h. Understanding its fields is central to writing correct Linux kernel PCI data structures-aware drivers. Here are the fields you will use constantly:

FieldMeaning
vendor / device16-bit IDs identifying the silicon manufacturer and the specific chip, read from configuration space
subsystem_vendor / subsystem_deviceBoard-level IDs, useful when several boards reuse the same chip
class24-bit class code; top byte is the base class (network, storage, display, and so on)
revisionSilicon revision, low byte of the class register
hdr_typeConfiguration header type — normal endpoint, PCI-to-PCI bridge, or CardBus bridge
bus / subordinateThe bus this function lives on, and (for bridges) the bus it forwards traffic to
driverPointer back to the bound pci_driver, once matched
devThe embedded struct device that ties this function into the generic driver model
irqAssigned interrupt line — meaning depends on INTx vs MSI/MSI-X, explained below
msi_enabled / msix_enabledSet once your driver successfully switches the device to message-signaled interrupts
dma_maskLargest DMA address the device can generate; used together with dma_set_mask()

The single most misunderstood field is irq. At boot, before any driver has touched interrupt configuration, irq simply holds the legacy INTx line the host controller pre-assigned. It stays that way until your driver explicitly calls pci_alloc_irq_vectors() (the modern replacement for the older pci_enable_msi() family). Once MSI mode is successfully negotiated, the kernel rewrites irq to the base Linux IRQ number of the allocated vector block — vector N then corresponds to pci_irq_vector(pdev, N), which is the API you should call rather than doing pointer arithmetic on pdev->irq yourself. This distinction matters a lot in real Linux kernel PCI data structures work, because reading irq too early, or assuming it never changes, is a recurring source of interrupt-routing bugs in embedded PCIe drivers.

Original Demo: Dumping pci_dev Fields From a Live Driver

Rather than reading about these fields in the abstract, let’s write a small original driver — ep_pci_dev_dump — that binds to a PCI function and logs the fields we just discussed straight into dmesg. We target the QEMU edu educational PCI device we introduced earlier in this series, since it is free, safe to experiment with, and needs no real hardware.

// ep_pci_dev_dump.c
#include <linux/module.h>
#include <linux/pci.h>
#include <linux/interrupt.h>

#define EP_VENDOR_ID 0x1234
#define EP_DEVICE_ID 0x11e8

static int ep_pci_dev_dump_probe(struct pci_dev *pdev,
                                  const struct pci_device_id *id)
{
    int nvec;

    if (pci_enable_device(pdev)) {
        dev_err(&pdev->dev, "ep_pci_dev_dump: enable failed\n");
        return -EIO;
    }

    dev_info(&pdev->dev, "ep_pci_dev_dump: vendor=0x%04x device=0x%04x\n",
              pdev->vendor, pdev->device);
    dev_info(&pdev->dev, "ep_pci_dev_dump: class=0x%06x revision=0x%02x\n",
              pdev->class, pdev->revision);
    dev_info(&pdev->dev, "ep_pci_dev_dump: hdr_type=%u irq(pre-MSI)=%u\n",
              pdev->hdr_type, pdev->irq);

    nvec = pci_alloc_irq_vectors(pdev, 1, 1, PCI_IRQ_MSI | PCI_IRQ_LEGACY);
    if (nvec < 0) {
        dev_err(&pdev->dev, "ep_pci_dev_dump: irq vector alloc failed\n");
        pci_disable_device(pdev);
        return nvec;
    }

    dev_info(&pdev->dev, "ep_pci_dev_dump: msi_enabled=%d final irq=%d\n",
              pdev->msi_enabled, pci_irq_vector(pdev, 0));

    return 0;
}

static void ep_pci_dev_dump_remove(struct pci_dev *pdev)
{
    pci_free_irq_vectors(pdev);
    pci_disable_device(pdev);
    dev_info(&pdev->dev, "ep_pci_dev_dump: removed\n");
}

static const struct pci_device_id ep_pci_dev_dump_ids[] = {
    { PCI_DEVICE(EP_VENDOR_ID, EP_DEVICE_ID) },
    { 0, }
};
MODULE_DEVICE_TABLE(pci, ep_pci_dev_dump_ids);

static struct pci_driver ep_pci_dev_dump_driver = {
    .name     = "ep_pci_dev_dump",
    .id_table = ep_pci_dev_dump_ids,
    .probe    = ep_pci_dev_dump_probe,
    .remove   = ep_pci_dev_dump_remove,
};

module_pci_driver(ep_pci_dev_dump_driver);
MODULE_LICENSE("GPL");
MODULE_DESCRIPTION("EmbeddedPathashala - dump struct pci_dev fields");

Build it with a minimal Kbuild file against your running kernel headers:

obj-m += ep_pci_dev_dump.o

all:
	make -C /lib/modules/$(shell uname -r)/build M=$(PWD) modules

clean:
	make -C /lib/modules/$(shell uname -r)/build M=$(PWD) clean
$ make
$ sudo insmod ep_pci_dev_dump.ko
$ dmesg | tail -n 6

Expected output on a QEMU guest booted with -device edu:

ep_pci_dev_dump: vendor=0x1234 device=0x11e8
ep_pci_dev_dump: class=0x00ff00 revision=0x10
ep_pci_dev_dump: hdr_type=0 irq(pre-MSI)=11
ep_pci_dev_dump: msi_enabled=1 final irq=27

Notice how irq jumps from the legacy pre-assigned value to a fresh MSI vector once pci_alloc_irq_vectors() succeeds — exactly the behavior we described above, now confirmed on a live system rather than taken on faith.

Legacy INTx vs MSI/MSI-X: A Quick Comparison

AspectLegacy INTxMSI / MSI-X
SignalingDedicated interrupt pin, shared across devicesIn-band memory write, one write per vector
SharingOften shared, requires driver to check status registersNot shared — each vector belongs to one function
Multiple vectorsNot possibleUp to thousands with MSI-X
Enable APIDefault, no call neededpci_alloc_irq_vectors(…, PCI_IRQ_MSI/MSIX)
Read vector IRQpdev->irq directlypci_irq_vector(pdev, index)

Real-World Use Cases

Understanding these Linux kernel PCI data structures is not academic — it directly shapes how you write drivers for NVMe SSDs, GPU and display controllers, Wi-Fi/Bluetooth combo chips on M.2 modules, and embedded PCIe accelerators found in automotive and industrial SoCs. Whenever you debug a driver that never receives interrupts, or a device that shows the wrong class code in lspci, you are debugging exactly the fields covered in this lecture.

Common Mistakes and Troubleshooting

  • Reading pdev->irq before calling pci_alloc_irq_vectors() and assuming it is the final MSI vector.
  • Forgetting to call pci_enable_device() before touching BARs or IRQs — the device stays inaccessible until this is done.
  • Not calling pci_free_irq_vectors() in remove(), leaking allocated MSI vectors on module unload.
  • Assuming every device supports MSI-X; always check the return value of pci_alloc_irq_vectors() and fall back to legacy INTx via PCI_IRQ_LEGACY.
  • Confusing class (24-bit, includes prog-if) with the base class byte alone when matching device types.

Best Practices

  • Always request MSI/MSI-X first and let pci_alloc_irq_vectors() fall back to legacy INTx automatically via the flags argument.
  • Use pci_irq_vector() rather than manual arithmetic on pdev->irq for every vector index.
  • Set an appropriate dma_mask early in probe() if your device performs DMA — this affects buffer allocation later in the driver’s life.
  • Release resources in the reverse order you acquired them: free IRQ vectors, then disable the device, mirroring the demo above.
  • For security, never trust configuration-space values from an unverified device blindly when building driver logic that affects system stability — validate ranges before using them as array indices or loop bounds.

Interview Questions

What creates struct pci_dev instances during boot?

The PCI core (drivers/pci/probe.c) creates them while walking the bus tree exposed by the host controller driver.

Why does pdev->irq change value after enabling MSI?

Because the kernel overwrites it with the base Linux IRQ number of the newly allocated MSI vector block.

What is the modern replacement for pci_enable_msi()?

pci_alloc_irq_vectors(), which can request MSI, MSI-X, or fall back to legacy INTx in one call.

What does the host controller layer provide to the PCI core?

Configuration-cycle access, memory/IO resource windows, and INTx/MSI interrupt line wiring for its bus.

How do you correctly read the IRQ for MSI vector 2?

Call pci_irq_vector(pdev, 2) instead of computing pdev->irq + 2 manually.

Summary and Key Takeaways

In this lecture of our free linux kernel development course we broke the Linux PCI software stack into its host controller, PCI core, and driver layers, then examined struct pci_dev field by field on a modern kernel, paying special attention to how the irq field behaves differently under legacy INTx and MSI/MSI-X. We closed with an original ep_pci_dev_dump driver that proves these behaviors on a live QEMU edu device rather than just describing them. Mastering these Linux kernel PCI data structures is a prerequisite for every advanced PCIe topic we cover later in this series, including DMA-capable endpoint drivers and PCIe power management.

Continue the Free Linux Kernel Development Course

Keep building real driver skills with EmbeddedPathashala’s free embedded systems course — every lecture uses original code, real dmesg output, and mainline kernel APIs.

Next Lecture Browse Full Course Index

Frequently Asked Questions

What is struct pci_dev used for in the Linux kernel?

It is the per-function object the PCI core creates for every detected PCI/PCIe device, holding its identification, configuration, and driver-binding state.

Where is struct pci_dev defined in kernel source?

In include/linux/pci.h, alongside struct pci_driver and struct pci_bus.

What is the difference between vendor/device and subsystem_vendor/subsystem_device?

vendor/device identify the silicon itself, while subsystem_vendor/subsystem_device identify the specific board or card that integrates that silicon.

Does every PCI device support MSI-X?

No. Support depends on the device’s capability list; always check the return value of pci_alloc_irq_vectors() and provide a legacy INTx fallback.

Why should I use pci_irq_vector() instead of pdev->irq directly?

Because pdev->irq only reliably represents the base vector; pci_irq_vector() correctly resolves any vector index for multi-vector MSI/MSI-X devices.

What kernel directory contains PCI host controller drivers?

drivers/pci/controller/, with shared IP-block drivers such as the DesignWare core grouped under drivers/pci/controller/dwc/.

Is this lecture part of a free Linux device drivers course?

Yes, this is part of EmbeddedPathashala’s free linux device drivers course covering the full PCI/PCIe subsystem from enumeration through DMA-capable endpoint drivers.

What is the class field in struct pci_dev used for?

It stores the 24-bit PCI class code, whose top byte identifies the device’s base class such as network, storage, or display controller.

PREV_LEC  |  NEXT_LEC

Leave a Reply

Your email address will not be published. Required fields are marked *