Linux Kernel Scatter-Gather DMA-Free Linux Device Drivers Course

PREV_LEC | NEXT_LEC

Linux Kernel Scatter-Gather DMA

A free Linux kernel development course lecture on building a real scatter-gather DMA mapping in a PCI device driver, using the modern dma_map_sgtable() API on a current kernel.

PCI / DMA Series
Lecture 15
Kernel 6.x APIs

This lecture is part of EmbeddedPathashala’s free Linux kernel development course, and it picks up right where the previous streaming DMA lecture left off. If you already understand dma_map_single() for one buffer at a time, this lecture extends that idea to scatter-gather DMA — mapping several physically scattered buffers in a single DMA operation using struct scatterlist and the modern dma_map_sgtable() API. We’ll build an original PCI driver, ep_pci_sgdma_demo, that allocates multiple kernel buffers, wires them into a scatter-gather table, maps them for a device, walks the mapped bus addresses, and cleanly tears everything down. This is core material for anyone following a free linux device drivers course and wanting to write real bus-mastering DMA drivers.

Topics Covered

struct scatterlist sg_alloc_table dma_map_sgtable for_each_sgtable_dma_sg sg_dma_address / sg_dma_len DMA sync for SG lists

What You Will Learn

  • Why a single DMA_TO_DEVICE / DMA_FROM_DEVICE mapping is not enough when a transfer touches many non-contiguous buffers
  • The layout of struct scatterlist and struct sg_table, and how the kernel stores per-entry bus addresses
  • The difference between sg_init_table() (static list) and sg_alloc_table() (dynamic allocation) for building a scatter-gather list
  • Why you should never poke the page field of a scatterlist entry directly, and why sg_set_page() / sg_set_buf() exist
  • How to map an entire list in one call with the modern dma_map_sgtable() wrapper, and why it has replaced the older raw dma_map_sg() call in new code
  • How to walk a mapped list safely with for_each_sgtable_dma_sg() and read each entry’s device-visible address with sg_dma_address() / sg_dma_len()
  • When you still need dma_sync_sg_for_cpu() / dma_sync_sg_for_device() between transfers
  • A complete, original ep_pci_sgdma_demo PCI driver you can build and load under QEMU

Prerequisites

This lecture builds directly on the earlier lectures in this free embedded linux course PCI series. Before continuing, you should already be comfortable with:

  • PCI driver registration and pci_enable_device() / pcim_enable_device() (covered in ldd2ch11_7)
  • DMA masks via dma_set_mask_and_coherent() (covered in ldd2ch11_13)
  • Single-buffer streaming mapping with dma_map_single() / dma_unmap_single() and the CPU/device ownership rule (covered in ldd2ch11_14)
  • Basic kernel module build flow (Makefile, insmod/rmmod, dmesg)

Why Scatter-Gather DMA Exists

Physical memory handed out by the kernel allocator is rarely one giant contiguous block once your buffers grow past a page or two. A network driver building a large outbound packet, a block driver assembling a multi-segment I/O request, or a video driver filling several frame buffers all end up with data spread across multiple, non-adjacent physical pages. Setting up a separate dma_map_single() call for every single page would work, but it means one hardware descriptor and one interrupt-safe bookkeeping entry per page — expensive in both CPU cycles and in device descriptor ring space.

Scatter-gather DMA solves this by letting the driver describe the entire set of buffers as one list, hand that whole list to the DMA mapping layer in a single call, and let capable hardware walk the list itself using its own descriptor chain. The CPU pays the mapping cost once; the device does the “gathering” of scattered source buffers (or “scattering” of one stream into multiple destination buffers) in hardware.

Scatter-Gather DMA Flow

Buffer A (page) Buffer B (page) Buffer C (partial) | | | v v v +———–+ +———–+ +————–+ | sg[0] | | sg[1] | | sg[2] | | page,off | | page,off | | page,off | | length | | length | | length | +———–+ +———–+ +————–+ \ | / \ | / \ | / dma_map_sgtable() | v device sees ONE descriptor list: [ dma_addr0,len0 ][ dma_addr1,len1 ][ dma_addr2,len2 ]

struct scatterlist and struct sg_table

Each buffer segment in the list is represented by one struct scatterlist entry. You never touch most of its fields directly — the kernel encodes the page pointer together with chaining flag bits inside page_link, so hand-editing it corrupts the list. Instead you always go through helper functions.

struct scatterlist {
    unsigned long   page_link;   /* encoded page + chain bit, never write directly */
    unsigned int    offset;      /* byte offset into the page */
    unsigned int    length;      /* length of this segment */
    dma_addr_t      dma_address; /* filled in by the mapping call */
#ifdef CONFIG_NEED_SG_DMA_LENGTH
    unsigned int    dma_length;  /* mapped length, may differ from length */
#endif
};

struct sg_table {
    struct scatterlist *sgl;     /* the array of entries */
    unsigned int        nents;   /* number of mapped entries (set by dma_map_sgtable) */
    unsigned int        orig_nents; /* number of entries you originally built */
};

Note the two counts inside sg_table. orig_nents is how many entries you allocated and filled in before mapping. nents is how many entries actually exist after mapping — because some IOMMU-backed platforms can merge two adjacent physical segments into a single device-visible bus address, the mapped count can be lower than the original count. Any code that loops over the mapped list must use nents, never orig_nents, which is exactly why the modern iteration macro (shown below) exists.

Building the List: sg_alloc_table() and sg_set_buf()

You have two ways to create the list. For a small, fixed number of entries known at compile time you can declare a struct scatterlist array on the stack or embedded in your device structure and call sg_init_table(). For a dynamic count decided at run time — the common case in a real driver — you allocate the array with sg_alloc_table(), which also takes care of chaining multiple pages of scatterlist entries together if the count is large.

struct sg_table sgt;
int ret;

ret = sg_alloc_table(&sgt, num_segments, GFP_KERNEL);
if (ret)
    return ret;

Once the table exists, fill each entry. If you already have a kernel virtual address (from kmalloc()), sg_set_buf() is the simplest helper — it derives the page and offset from the pointer for you:

sg_set_buf(&sgt.sgl[i], kernel_ptr, seg_len);

If instead you already hold struct page * pointers (common when working with user pages pinned via pin_user_pages(), or with pages you allocated yourself), use sg_set_page() and walk the list with for_each_sg():

struct scatterlist *sg;
int i;

for_each_sg(sgt.sgl, sg, sgt.orig_nents, i)
    sg_set_page(sg, pages[i], seg_len[i], seg_off[i]);

Mapping the List: dma_map_sgtable()

Older drivers call dma_map_sg(dev, sgt.sgl, sgt.orig_nents, dir) directly, which returns the mapped entry count as a plain integer that you then have to remember to store back into nents yourself. It’s easy to forget, and easy to unmap with the wrong count. The kernel’s DMA API now provides dma_map_sgtable(), which operates on the whole struct sg_table and updates nents internally, so the object stays self-consistent:

ret = dma_map_sgtable(&pdev->dev, &sgt, DMA_TO_DEVICE, 0);
if (ret) {
    dev_err(&pdev->dev, "sg map failed: %d\n", ret);
    sg_free_table(&sgt);
    return ret;
}

The direction argument follows the same DMA_TO_DEVICE / DMA_FROM_DEVICE / DMA_BIDIRECTIONAL flags used for single-buffer streaming mappings. The final argument is an attrs bitmask — pass 0 for the common case.

Walking the Mapped List

After a successful map, you hand the device its descriptor addresses by iterating the mapped entries — using sgt.nents, not the original count. The macro to use for this is for_each_sgtable_dma_sg(), which is IOMMU-merge aware:

struct scatterlist *sg;
int i;

for_each_sgtable_dma_sg(&sgt, sg, i) {
    dma_addr_t bus_addr = sg_dma_address(sg);
    unsigned int bus_len = sg_dma_len(sg);

    dev_info(&pdev->dev, "sg[%d]: dma_addr=%pad len=%u\n",
              i, &bus_addr, bus_len);

    /* program one hardware descriptor here, e.g.:
     * write_descriptor(dev, bus_addr, bus_len);
     */
}

sg_dma_address() and sg_dma_len() are architecture-portable accessors — the actual storage location of the bus address inside the scatterlist entry differs between architectures, which is exactly why you should never read dma_address as a plain struct field in portable code.

Syncing and Unmapping

If the CPU needs to touch the buffers again between DMA transfers without a full unmap/remap cycle, use the scatter-gather sync calls, mirroring the single-buffer versions from the previous lecture:

dma_sync_sgtable_for_cpu(&pdev->dev, &sgt, DMA_FROM_DEVICE);
/* CPU reads/writes buffers safely here */
dma_sync_sgtable_for_device(&pdev->dev, &sgt, DMA_FROM_DEVICE);

When the transfer is completely done, unmap and free the table in that order:

dma_unmap_sgtable(&pdev->dev, &sgt, DMA_TO_DEVICE, 0);
sg_free_table(&sgt);

Complete Driver: ep_pci_sgdma_demo

The following original demo driver builds on the QEMU edu educational PCI device used throughout this series. It allocates three separate kernel buffers of different sizes, assembles them into one scatter-gather table, maps the table, logs the bus address of every mapped segment, and tears the mapping down on module removal. The point of this driver is to make the mapping/walking/unmapping lifecycle completely visible in dmesg, which is the part most tutorials skip.

#include <linux/module.h>
#include <linux/pci.h>
#include <linux/dma-mapping.h>
#include <linux/slab.h>
#include <linux/scatterlist.h>

#define EP_SG_SEGMENTS   3
#define EP_SEG_SIZE      PAGE_SIZE

struct ep_sgdma_priv {
    struct pci_dev  *pdev;
    struct sg_table  sgt;
    void            *buf[EP_SG_SEGMENTS];
};

static int ep_sgdma_build_and_map(struct ep_sgdma_priv *priv)
{
    struct scatterlist *sg;
    int i, ret;

    ret = sg_alloc_table(&priv->sgt, EP_SG_SEGMENTS, GFP_KERNEL);
    if (ret)
        return ret;

    for_each_sg(priv->sgt.sgl, sg, EP_SG_SEGMENTS, i) {
        priv->buf[i] = kmalloc(EP_SEG_SIZE, GFP_KERNEL | GFP_DMA);
        if (!priv->buf[i]) {
            ret = -ENOMEM;
            goto err_free_bufs;
        }
        memset(priv->buf[i], 0xA0 + i, EP_SEG_SIZE);
        sg_set_buf(sg, priv->buf[i], EP_SEG_SIZE);
    }

    ret = dma_map_sgtable(&priv->pdev->dev, &priv->sgt, DMA_TO_DEVICE, 0);
    if (ret) {
        dev_err(&priv->pdev->dev, "dma_map_sgtable failed: %d\n", ret);
        goto err_free_bufs;
    }

    for_each_sgtable_dma_sg(&priv->sgt, sg, i) {
        dma_addr_t addr = sg_dma_address(sg);
        unsigned int len = sg_dma_len(sg);

        dev_info(&priv->pdev->dev,
                 "ep_sgdma: segment %d -> dma_addr=%pad len=%u\n",
                 i, &addr, len);
    }

    return 0;

err_free_bufs:
    while (--i >= 0)
        kfree(priv->buf[i]);
    sg_free_table(&priv->sgt);
    return ret;
}

static void ep_sgdma_unmap_and_free(struct ep_sgdma_priv *priv)
{
    int i;

    dma_unmap_sgtable(&priv->pdev->dev, &priv->sgt, DMA_TO_DEVICE, 0);
    sg_free_table(&priv->sgt);

    for (i = 0; i buf[i]);
}

static int ep_pci_sgdma_probe(struct pci_dev *pdev,
                               const struct pci_device_id *id)
{
    struct ep_sgdma_priv *priv;
    int ret;

    priv = devm_kzalloc(&pdev->dev, sizeof(*priv), GFP_KERNEL);
    if (!priv)
        return -ENOMEM;
    priv->pdev = pdev;

    ret = pcim_enable_device(pdev);
    if (ret)
        return ret;
    pci_set_master(pdev);

    ret = dma_set_mask_and_coherent(&pdev->dev, DMA_BIT_MASK(28));
    if (ret) {
        dev_err(&pdev->dev, "no usable DMA mask\n");
        return ret;
    }

    pci_set_drvdata(pdev, priv);

    ret = ep_sgdma_build_and_map(priv);
    if (ret)
        return ret;

    dev_info(&pdev->dev, "ep_pci_sgdma_demo: probe complete, %u segments mapped\n",
             priv->sgt.nents);
    return 0;
}

static void ep_pci_sgdma_remove(struct pci_dev *pdev)
{
    struct ep_sgdma_priv *priv = pci_get_drvdata(pdev);

    ep_sgdma_unmap_and_free(priv);
    dev_info(&pdev->dev, "ep_pci_sgdma_demo: unmapped and cleaned up\n");
}

static const struct pci_device_id ep_sgdma_ids[] = {
    { PCI_DEVICE(0x1234, 0x11e8) }, /* QEMU edu device */
    { }
};
MODULE_DEVICE_TABLE(pci, ep_sgdma_ids);

static struct pci_driver ep_pci_sgdma_driver = {
    .name     = "ep_pci_sgdma_demo",
    .id_table = ep_sgdma_ids,
    .probe    = ep_pci_sgdma_probe,
    .remove   = ep_pci_sgdma_remove,
};
module_pci_driver(ep_pci_sgdma_driver);

MODULE_LICENSE("GPL");
MODULE_DESCRIPTION("EmbeddedPathashala scatter-gather DMA demo for the free Linux kernel development course");

Build and Run Walkthrough

Build the module against your running kernel’s headers, exactly as in earlier lectures of this free linux development course:

make -C /lib/modules/$(uname -r)/build M=$(pwd) modules

Boot QEMU with the edu device attached (same flag used across this series), then load the module:

qemu-system-x86_64 -device edu -kernel bzImage -append "console=ttyS0" ...
sudo insmod ep_pci_sgdma_demo.ko
dmesg | tail -n 6

Expected output on probe — three distinct device-visible bus addresses, one per segment:

[  12.104332] ep_pci_sgdma_demo 0000:00:03.0: ep_sgdma: segment 0 -> dma_addr=0x00000000ff01000 len=4096
[  12.104351] ep_pci_sgdma_demo 0000:00:03.0: ep_sgdma: segment 1 -> dma_addr=0x00000000ff02000 len=4096
[  12.104367] ep_pci_sgdma_demo 0000:00:03.0: ep_sgdma: segment 2 -> dma_addr=0x00000000ff03000 len=4096
[  12.104380] ep_pci_sgdma_demo 0000:00:03.0: ep_pci_sgdma_demo: probe complete, 3 segments mapped

On rmmod, the driver unmaps and frees cleanly:

sudo rmmod ep_pci_sgdma_demo
dmesg | tail -n 2
[  40.552110] ep_pci_sgdma_demo 0000:00:03.0: ep_pci_sgdma_demo: unmapped and cleaned up

Common Mistakes and Troubleshooting

  • Unmapping with orig_nents instead of nents: if an IOMMU merged segments during mapping, unmapping with the pre-map count corrupts IOMMU state. dma_unmap_sgtable() avoids this by reading nents from the table itself.
  • Writing to sg->page_link or dma_address directly: always go through sg_set_page()/sg_set_buf() and sg_dma_address()/sg_dma_len() — the internal encoding is not a plain pointer.
  • Forgetting sg_free_table(): unmapping releases the DMA/IOMMU mapping but not the scatterlist array itself; both calls are required.
  • Touching CPU buffers after mapping without a sync call: once mapped for DMA_TO_DEVICE/DMA_FROM_DEVICE, the buffer is owned by the device until you sync it back, exactly as with single-buffer streaming mappings.
  • Ignoring the return value of dma_map_sgtable(): a mapping failure (for example, IOMMU exhaustion) must be handled — never assume the call always succeeds.

Best Practices

  • Prefer dma_map_sgtable() / dma_unmap_sgtable() over the raw dma_map_sg() / dma_unmap_sg() pair in all new driver code — the table-based API keeps entry counts consistent automatically.
  • Always iterate the mapped list with for_each_sgtable_dma_sg(), never a hand-rolled loop over the raw array, so IOMMU-merged entries are handled correctly.
  • Keep the number of segments as small as your hardware descriptor ring realistically allows — every segment still costs one descriptor slot even though the mapping call itself is batched.
  • For performance-sensitive paths, avoid rebuilding the scatter-gather table on every transfer; reuse it and rely on the sync calls where the buffer set doesn’t change between transfers.
  • For security, never trust a scatter-gather length or offset that originated from user space without validating it against the actual allocated buffer size before calling sg_set_buf() or sg_set_page().

Summary and Key Takeaways

Scatter-gather DMA lets a driver describe many physically scattered buffers as one logical transfer instead of issuing one single-buffer mapping per segment. struct sg_table tracks both the original segment count and the post-mapping count, and the modern dma_map_sgtable() / dma_unmap_sgtable() pair keeps those two numbers consistent automatically, which is why they’ve replaced the older raw dma_map_sg() call in current kernel code. Combined with for_each_sgtable_dma_sg(), sg_dma_address(), and sg_dma_len(), this gives a portable, IOMMU-safe way to hand a device a full descriptor list in one shot. This closes out the DMA portion of the PCI driver series in this free linux kernel development course — the next lecture moves on to the NVMEM framework.

Frequently Asked Questions

What is the difference between dma_map_sg() and dma_map_sgtable()?

dma_map_sg() returns the mapped entry count as a plain integer that the driver must store back manually. dma_map_sgtable() operates on the whole struct sg_table and updates its nents field internally, keeping the original and mapped counts self-consistent — this is why it’s the recommended API for new drivers.

Why can the mapped entry count be lower than the number of buffers I allocated?

On platforms with an IOMMU, physically adjacent segments can be merged into a single device-visible bus address during mapping, which reduces the effective entry count. This is exactly why iteration must use sgt.nents, not the original allocation count.

Can I write directly to a scatterlist entry’s dma_address field?

No. Always use the sg_dma_address() and sg_dma_len() accessor macros. The exact storage location of these values inside the structure is architecture-dependent, so reading or writing the field directly breaks portability.

Do I still need sg_free_table() after calling dma_unmap_sgtable()?

Yes. Unmapping releases the DMA/IOMMU mapping only. sg_free_table() releases the scatterlist array itself (and any chained pages it required), and both calls are needed for a clean teardown.

When should I use sg_set_buf() versus sg_set_page()?

Use sg_set_buf() when you already have a kernel virtual address, such as from kmalloc() — it derives the page and offset for you. Use sg_set_page() when you already hold struct page * pointers directly, which is common when working with pinned user pages.

Is scatter-gather DMA only useful for large transfers?

It’s most valuable whenever data spans multiple physically non-contiguous buffers, regardless of total size — network packet fragments, multi-segment block I/O requests, and video frame buffers all commonly use it even for moderate transfer sizes.

What happens if dma_map_sgtable() fails partway through mapping?

It returns a non-zero error code and does not leave a partially-mapped table behind — treat any non-zero return as a full failure, free the scatterlist table, and release the buffers you allocated before returning the error up the probe path.

Does this lecture belong to a free course?

Yes. This lecture is part of EmbeddedPathashala’s free Linux kernel development course, which also covers character devices, platform drivers, DMA, interrupts, and PCI/PCIe as part of a broader free embedded systems course and free Linux device drivers course.

Continue the Free Linux Kernel Development Course

Next up: the NVMEM framework for non-volatile storage devices like EEPROM.

Next Lecture Browse Full Course Index

PREV_LEC | NEXT_LEC

Leave a Reply

Your email address will not be published. Required fields are marked *