If you are following our free Linux kernel development course, you already know how a PCI device gets discovered on the bus. In this lecture we go one level deeper into the Linux kernel PCI data structures that every PCI/PCIe driver author must understand before writing a single line of probe() code. We will walk through the three software layers that bring a PCI device to life inside the kernel, then dissect struct pci_dev field by field using the current mainline kernel source, and finish by writing an original driver that dumps this information live from dmesg. This is exactly the kind of practical, from-scratch content you’d expect from a free linux device drivers course — no copy-pasted book listings, only concepts explained in plain language and code you can build yourself.
What You Will Learn
- The three software layers of the Linux PCI stack
- Every important field of struct pci_dev on kernel 6.x
- How the irq field changes meaning under legacy INTx vs MSI/MSI-X
- Writing an original driver that reads and prints pci_dev fields
- Common mistakes when accessing PCI device state
Prerequisites
This lecture continues our PCI/PCIe device driver series. Before reading further, we recommend finishing the earlier lectures on PCI/PCIe bus topology and enumeration, header types and BDF addressing, PCI address spaces and BAR mapping, and PCI interrupt distribution (legacy INTx vs MSI/MSI-X). A working knowledge of C, basic kernel module loading (insmod/rmmod), and a QEMU test environment is assumed — all covered earlier in this free linux kernel development course.
The Three Layers of the Linux PCI Software Stack
One of the most confusing parts for newcomers to Linux kernel PCI data structures is understanding *who* actually creates the device objects the driver eventually binds to. The kernel splits this responsibility across three cooperating layers, each with a distinct job:
Architecture PCI glue is the earliest piece to run. On modern ARM64 and most embedded SoCs this work has largely moved into the generic PCI core plus the host controller driver, but conceptually its job stays the same: get the bus numbered and ready before anything else touches it.
The host controller (often called the root complex) is SoC-specific. You will find these drivers under drivers/pci/controller/ in a modern kernel tree — for example the generic DesignWare-based controllers live under drivers/pci/controller/dwc/, while other SoC vendors ship their own controller driver in the same directory. This layer is responsible for generating configuration-space read/write cycles, exposing the memory/IO resource windows the SoC provides, and wiring up the INTx and MSI interrupt lines. Without a correctly initialized host controller, the PCI core has no bus to scan at all.
The PCI core, implemented mainly in drivers/pci/probe.c, is where the actual struct pci_dev and struct pci_bus objects are born. The core walks every bus/device/function combination reachable through the host controller, reads the standard configuration header from each function that responds, and allocates a pci_dev instance for it. It also creates the sysfs entries under /sys/bus/pci/devices/ that you use every day with tools like lspci, handles PCI Express port/bridge devices, and provides the MSI/MSI-X interrupt allocation framework your driver will call into.
Once the core has built this tree, it matches each pci_dev against the pci_device_id tables registered by loaded pci_driver structures and calls your probe() function — which brings us to the object your driver actually works with: struct pci_dev.
Inside struct pci_dev on a Modern Kernel
Every PCI or PCIe function detected on the bus gets exactly one struct pci_dev, defined in include/linux/pci.h. Understanding its fields is central to writing correct Linux kernel PCI data structures-aware drivers. Here are the fields you will use constantly:
| Field | Meaning |
|---|---|
| vendor / device | 16-bit IDs identifying the silicon manufacturer and the specific chip, read from configuration space |
| subsystem_vendor / subsystem_device | Board-level IDs, useful when several boards reuse the same chip |
| class | 24-bit class code; top byte is the base class (network, storage, display, and so on) |
| revision | Silicon revision, low byte of the class register |
| hdr_type | Configuration header type — normal endpoint, PCI-to-PCI bridge, or CardBus bridge |
| bus / subordinate | The bus this function lives on, and (for bridges) the bus it forwards traffic to |
| driver | Pointer back to the bound pci_driver, once matched |
| dev | The embedded struct device that ties this function into the generic driver model |
| irq | Assigned interrupt line — meaning depends on INTx vs MSI/MSI-X, explained below |
| msi_enabled / msix_enabled | Set once your driver successfully switches the device to message-signaled interrupts |
| dma_mask | Largest DMA address the device can generate; used together with dma_set_mask() |
The single most misunderstood field is irq. At boot, before any driver has touched interrupt configuration, irq simply holds the legacy INTx line the host controller pre-assigned. It stays that way until your driver explicitly calls pci_alloc_irq_vectors() (the modern replacement for the older pci_enable_msi() family). Once MSI mode is successfully negotiated, the kernel rewrites irq to the base Linux IRQ number of the allocated vector block — vector N then corresponds to pci_irq_vector(pdev, N), which is the API you should call rather than doing pointer arithmetic on pdev->irq yourself. This distinction matters a lot in real Linux kernel PCI data structures work, because reading irq too early, or assuming it never changes, is a recurring source of interrupt-routing bugs in embedded PCIe drivers.
Original Demo: Dumping pci_dev Fields From a Live Driver
Rather than reading about these fields in the abstract, let’s write a small original driver — ep_pci_dev_dump — that binds to a PCI function and logs the fields we just discussed straight into dmesg. We target the QEMU edu educational PCI device we introduced earlier in this series, since it is free, safe to experiment with, and needs no real hardware.
// ep_pci_dev_dump.c
#include <linux/module.h>
#include <linux/pci.h>
#include <linux/interrupt.h>
#define EP_VENDOR_ID 0x1234
#define EP_DEVICE_ID 0x11e8
static int ep_pci_dev_dump_probe(struct pci_dev *pdev,
const struct pci_device_id *id)
{
int nvec;
if (pci_enable_device(pdev)) {
dev_err(&pdev->dev, "ep_pci_dev_dump: enable failed\n");
return -EIO;
}
dev_info(&pdev->dev, "ep_pci_dev_dump: vendor=0x%04x device=0x%04x\n",
pdev->vendor, pdev->device);
dev_info(&pdev->dev, "ep_pci_dev_dump: class=0x%06x revision=0x%02x\n",
pdev->class, pdev->revision);
dev_info(&pdev->dev, "ep_pci_dev_dump: hdr_type=%u irq(pre-MSI)=%u\n",
pdev->hdr_type, pdev->irq);
nvec = pci_alloc_irq_vectors(pdev, 1, 1, PCI_IRQ_MSI | PCI_IRQ_LEGACY);
if (nvec < 0) {
dev_err(&pdev->dev, "ep_pci_dev_dump: irq vector alloc failed\n");
pci_disable_device(pdev);
return nvec;
}
dev_info(&pdev->dev, "ep_pci_dev_dump: msi_enabled=%d final irq=%d\n",
pdev->msi_enabled, pci_irq_vector(pdev, 0));
return 0;
}
static void ep_pci_dev_dump_remove(struct pci_dev *pdev)
{
pci_free_irq_vectors(pdev);
pci_disable_device(pdev);
dev_info(&pdev->dev, "ep_pci_dev_dump: removed\n");
}
static const struct pci_device_id ep_pci_dev_dump_ids[] = {
{ PCI_DEVICE(EP_VENDOR_ID, EP_DEVICE_ID) },
{ 0, }
};
MODULE_DEVICE_TABLE(pci, ep_pci_dev_dump_ids);
static struct pci_driver ep_pci_dev_dump_driver = {
.name = "ep_pci_dev_dump",
.id_table = ep_pci_dev_dump_ids,
.probe = ep_pci_dev_dump_probe,
.remove = ep_pci_dev_dump_remove,
};
module_pci_driver(ep_pci_dev_dump_driver);
MODULE_LICENSE("GPL");
MODULE_DESCRIPTION("EmbeddedPathashala - dump struct pci_dev fields");
Build it with a minimal Kbuild file against your running kernel headers:
obj-m += ep_pci_dev_dump.o
all:
make -C /lib/modules/$(shell uname -r)/build M=$(PWD) modules
clean:
make -C /lib/modules/$(shell uname -r)/build M=$(PWD) clean
$ make
$ sudo insmod ep_pci_dev_dump.ko
$ dmesg | tail -n 6
Expected output on a QEMU guest booted with -device edu:
ep_pci_dev_dump: vendor=0x1234 device=0x11e8
ep_pci_dev_dump: class=0x00ff00 revision=0x10
ep_pci_dev_dump: hdr_type=0 irq(pre-MSI)=11
ep_pci_dev_dump: msi_enabled=1 final irq=27
Notice how irq jumps from the legacy pre-assigned value to a fresh MSI vector once pci_alloc_irq_vectors() succeeds — exactly the behavior we described above, now confirmed on a live system rather than taken on faith.
Legacy INTx vs MSI/MSI-X: A Quick Comparison
| Aspect | Legacy INTx | MSI / MSI-X |
|---|---|---|
| Signaling | Dedicated interrupt pin, shared across devices | In-band memory write, one write per vector |
| Sharing | Often shared, requires driver to check status registers | Not shared — each vector belongs to one function |
| Multiple vectors | Not possible | Up to thousands with MSI-X |
| Enable API | Default, no call needed | pci_alloc_irq_vectors(…, PCI_IRQ_MSI/MSIX) |
| Read vector IRQ | pdev->irq directly | pci_irq_vector(pdev, index) |
Real-World Use Cases
Understanding these Linux kernel PCI data structures is not academic — it directly shapes how you write drivers for NVMe SSDs, GPU and display controllers, Wi-Fi/Bluetooth combo chips on M.2 modules, and embedded PCIe accelerators found in automotive and industrial SoCs. Whenever you debug a driver that never receives interrupts, or a device that shows the wrong class code in lspci, you are debugging exactly the fields covered in this lecture.
Common Mistakes and Troubleshooting
- Reading
pdev->irqbefore callingpci_alloc_irq_vectors()and assuming it is the final MSI vector. - Forgetting to call
pci_enable_device()before touching BARs or IRQs — the device stays inaccessible until this is done. - Not calling
pci_free_irq_vectors()in remove(), leaking allocated MSI vectors on module unload. - Assuming every device supports MSI-X; always check the return value of
pci_alloc_irq_vectors()and fall back to legacy INTx viaPCI_IRQ_LEGACY. - Confusing
class(24-bit, includes prog-if) with the base class byte alone when matching device types.
Best Practices
- Always request MSI/MSI-X first and let
pci_alloc_irq_vectors()fall back to legacy INTx automatically via the flags argument. - Use
pci_irq_vector()rather than manual arithmetic onpdev->irqfor every vector index. - Set an appropriate
dma_maskearly in probe() if your device performs DMA — this affects buffer allocation later in the driver’s life. - Release resources in the reverse order you acquired them: free IRQ vectors, then disable the device, mirroring the demo above.
- For security, never trust configuration-space values from an unverified device blindly when building driver logic that affects system stability — validate ranges before using them as array indices or loop bounds.
Interview Questions
What creates struct pci_dev instances during boot?
The PCI core (drivers/pci/probe.c) creates them while walking the bus tree exposed by the host controller driver.
Why does pdev->irq change value after enabling MSI?
Because the kernel overwrites it with the base Linux IRQ number of the newly allocated MSI vector block.
What is the modern replacement for pci_enable_msi()?
pci_alloc_irq_vectors(), which can request MSI, MSI-X, or fall back to legacy INTx in one call.
What does the host controller layer provide to the PCI core?
Configuration-cycle access, memory/IO resource windows, and INTx/MSI interrupt line wiring for its bus.
How do you correctly read the IRQ for MSI vector 2?
Call pci_irq_vector(pdev, 2) instead of computing pdev->irq + 2 manually.
Summary and Key Takeaways
In this lecture of our free linux kernel development course we broke the Linux PCI software stack into its host controller, PCI core, and driver layers, then examined struct pci_dev field by field on a modern kernel, paying special attention to how the irq field behaves differently under legacy INTx and MSI/MSI-X. We closed with an original ep_pci_dev_dump driver that proves these behaviors on a live QEMU edu device rather than just describing them. Mastering these Linux kernel PCI data structures is a prerequisite for every advanced PCIe topic we cover later in this series, including DMA-capable endpoint drivers and PCIe power management.
Continue the Free Linux Kernel Development Course
Keep building real driver skills with EmbeddedPathashala’s free embedded systems course — every lecture uses original code, real dmesg output, and mainline kernel APIs.
Next Lecture Browse Full Course IndexFrequently Asked Questions
What is struct pci_dev used for in the Linux kernel?
It is the per-function object the PCI core creates for every detected PCI/PCIe device, holding its identification, configuration, and driver-binding state.
Where is struct pci_dev defined in kernel source?
In include/linux/pci.h, alongside struct pci_driver and struct pci_bus.
What is the difference between vendor/device and subsystem_vendor/subsystem_device?
vendor/device identify the silicon itself, while subsystem_vendor/subsystem_device identify the specific board or card that integrates that silicon.
Does every PCI device support MSI-X?
No. Support depends on the device’s capability list; always check the return value of pci_alloc_irq_vectors() and provide a legacy INTx fallback.
Why should I use pci_irq_vector() instead of pdev->irq directly?
Because pdev->irq only reliably represents the base vector; pci_irq_vector() correctly resolves any vector index for multi-vector MSI/MSI-X devices.
What kernel directory contains PCI host controller drivers?
drivers/pci/controller/, with shared IP-block drivers such as the DesignWare core grouped under drivers/pci/controller/dwc/.
Is this lecture part of a free Linux device drivers course?
Yes, this is part of EmbeddedPathashala’s free linux device drivers course covering the full PCI/PCIe subsystem from enumeration through DMA-capable endpoint drivers.
What is the class field in struct pci_dev used for?
It stores the 24-bit PCI class code, whose top byte identifies the device’s base class such as network, storage, or display controller.
