Linux Kernel Scatter-Gather DMA
A free Linux kernel development course lecture on building a real scatter-gather DMA mapping in a PCI device driver, using the modern dma_map_sgtable() API on a current kernel.
This lecture is part of EmbeddedPathashala’s free Linux kernel development course, and it picks up right where the previous streaming DMA lecture left off. If you already understand dma_map_single() for one buffer at a time, this lecture extends that idea to scatter-gather DMA — mapping several physically scattered buffers in a single DMA operation using struct scatterlist and the modern dma_map_sgtable() API. We’ll build an original PCI driver, ep_pci_sgdma_demo, that allocates multiple kernel buffers, wires them into a scatter-gather table, maps them for a device, walks the mapped bus addresses, and cleanly tears everything down. This is core material for anyone following a free linux device drivers course and wanting to write real bus-mastering DMA drivers.
Topics Covered
What You Will Learn
- Why a single DMA_TO_DEVICE / DMA_FROM_DEVICE mapping is not enough when a transfer touches many non-contiguous buffers
- The layout of
struct scatterlistandstruct sg_table, and how the kernel stores per-entry bus addresses - The difference between
sg_init_table()(static list) andsg_alloc_table()(dynamic allocation) for building a scatter-gather list - Why you should never poke the
pagefield of a scatterlist entry directly, and whysg_set_page()/sg_set_buf()exist - How to map an entire list in one call with the modern
dma_map_sgtable()wrapper, and why it has replaced the older rawdma_map_sg()call in new code - How to walk a mapped list safely with
for_each_sgtable_dma_sg()and read each entry’s device-visible address withsg_dma_address()/sg_dma_len() - When you still need
dma_sync_sg_for_cpu()/dma_sync_sg_for_device()between transfers - A complete, original
ep_pci_sgdma_demoPCI driver you can build and load under QEMU
Prerequisites
This lecture builds directly on the earlier lectures in this free embedded linux course PCI series. Before continuing, you should already be comfortable with:
- PCI driver registration and
pci_enable_device()/pcim_enable_device()(covered in ldd2ch11_7) - DMA masks via
dma_set_mask_and_coherent()(covered in ldd2ch11_13) - Single-buffer streaming mapping with
dma_map_single()/dma_unmap_single()and the CPU/device ownership rule (covered in ldd2ch11_14) - Basic kernel module build flow (Makefile,
insmod/rmmod,dmesg)
Why Scatter-Gather DMA Exists
Physical memory handed out by the kernel allocator is rarely one giant contiguous block once your buffers grow past a page or two. A network driver building a large outbound packet, a block driver assembling a multi-segment I/O request, or a video driver filling several frame buffers all end up with data spread across multiple, non-adjacent physical pages. Setting up a separate dma_map_single() call for every single page would work, but it means one hardware descriptor and one interrupt-safe bookkeeping entry per page — expensive in both CPU cycles and in device descriptor ring space.
Scatter-gather DMA solves this by letting the driver describe the entire set of buffers as one list, hand that whole list to the DMA mapping layer in a single call, and let capable hardware walk the list itself using its own descriptor chain. The CPU pays the mapping cost once; the device does the “gathering” of scattered source buffers (or “scattering” of one stream into multiple destination buffers) in hardware.
Scatter-Gather DMA Flow
struct scatterlist and struct sg_table
Each buffer segment in the list is represented by one struct scatterlist entry. You never touch most of its fields directly — the kernel encodes the page pointer together with chaining flag bits inside page_link, so hand-editing it corrupts the list. Instead you always go through helper functions.
struct scatterlist {
unsigned long page_link; /* encoded page + chain bit, never write directly */
unsigned int offset; /* byte offset into the page */
unsigned int length; /* length of this segment */
dma_addr_t dma_address; /* filled in by the mapping call */
#ifdef CONFIG_NEED_SG_DMA_LENGTH
unsigned int dma_length; /* mapped length, may differ from length */
#endif
};
struct sg_table {
struct scatterlist *sgl; /* the array of entries */
unsigned int nents; /* number of mapped entries (set by dma_map_sgtable) */
unsigned int orig_nents; /* number of entries you originally built */
};
Note the two counts inside sg_table. orig_nents is how many entries you allocated and filled in before mapping. nents is how many entries actually exist after mapping — because some IOMMU-backed platforms can merge two adjacent physical segments into a single device-visible bus address, the mapped count can be lower than the original count. Any code that loops over the mapped list must use nents, never orig_nents, which is exactly why the modern iteration macro (shown below) exists.
Building the List: sg_alloc_table() and sg_set_buf()
You have two ways to create the list. For a small, fixed number of entries known at compile time you can declare a struct scatterlist array on the stack or embedded in your device structure and call sg_init_table(). For a dynamic count decided at run time — the common case in a real driver — you allocate the array with sg_alloc_table(), which also takes care of chaining multiple pages of scatterlist entries together if the count is large.
struct sg_table sgt;
int ret;
ret = sg_alloc_table(&sgt, num_segments, GFP_KERNEL);
if (ret)
return ret;
Once the table exists, fill each entry. If you already have a kernel virtual address (from kmalloc()), sg_set_buf() is the simplest helper — it derives the page and offset from the pointer for you:
sg_set_buf(&sgt.sgl[i], kernel_ptr, seg_len);
If instead you already hold struct page * pointers (common when working with user pages pinned via pin_user_pages(), or with pages you allocated yourself), use sg_set_page() and walk the list with for_each_sg():
struct scatterlist *sg;
int i;
for_each_sg(sgt.sgl, sg, sgt.orig_nents, i)
sg_set_page(sg, pages[i], seg_len[i], seg_off[i]);
Mapping the List: dma_map_sgtable()
Older drivers call dma_map_sg(dev, sgt.sgl, sgt.orig_nents, dir) directly, which returns the mapped entry count as a plain integer that you then have to remember to store back into nents yourself. It’s easy to forget, and easy to unmap with the wrong count. The kernel’s DMA API now provides dma_map_sgtable(), which operates on the whole struct sg_table and updates nents internally, so the object stays self-consistent:
ret = dma_map_sgtable(&pdev->dev, &sgt, DMA_TO_DEVICE, 0);
if (ret) {
dev_err(&pdev->dev, "sg map failed: %d\n", ret);
sg_free_table(&sgt);
return ret;
}
The direction argument follows the same DMA_TO_DEVICE / DMA_FROM_DEVICE / DMA_BIDIRECTIONAL flags used for single-buffer streaming mappings. The final argument is an attrs bitmask — pass 0 for the common case.
Walking the Mapped List
After a successful map, you hand the device its descriptor addresses by iterating the mapped entries — using sgt.nents, not the original count. The macro to use for this is for_each_sgtable_dma_sg(), which is IOMMU-merge aware:
struct scatterlist *sg;
int i;
for_each_sgtable_dma_sg(&sgt, sg, i) {
dma_addr_t bus_addr = sg_dma_address(sg);
unsigned int bus_len = sg_dma_len(sg);
dev_info(&pdev->dev, "sg[%d]: dma_addr=%pad len=%u\n",
i, &bus_addr, bus_len);
/* program one hardware descriptor here, e.g.:
* write_descriptor(dev, bus_addr, bus_len);
*/
}
sg_dma_address() and sg_dma_len() are architecture-portable accessors — the actual storage location of the bus address inside the scatterlist entry differs between architectures, which is exactly why you should never read dma_address as a plain struct field in portable code.
Syncing and Unmapping
If the CPU needs to touch the buffers again between DMA transfers without a full unmap/remap cycle, use the scatter-gather sync calls, mirroring the single-buffer versions from the previous lecture:
dma_sync_sgtable_for_cpu(&pdev->dev, &sgt, DMA_FROM_DEVICE);
/* CPU reads/writes buffers safely here */
dma_sync_sgtable_for_device(&pdev->dev, &sgt, DMA_FROM_DEVICE);
When the transfer is completely done, unmap and free the table in that order:
dma_unmap_sgtable(&pdev->dev, &sgt, DMA_TO_DEVICE, 0);
sg_free_table(&sgt);
Complete Driver: ep_pci_sgdma_demo
The following original demo driver builds on the QEMU edu educational PCI device used throughout this series. It allocates three separate kernel buffers of different sizes, assembles them into one scatter-gather table, maps the table, logs the bus address of every mapped segment, and tears the mapping down on module removal. The point of this driver is to make the mapping/walking/unmapping lifecycle completely visible in dmesg, which is the part most tutorials skip.
#include <linux/module.h>
#include <linux/pci.h>
#include <linux/dma-mapping.h>
#include <linux/slab.h>
#include <linux/scatterlist.h>
#define EP_SG_SEGMENTS 3
#define EP_SEG_SIZE PAGE_SIZE
struct ep_sgdma_priv {
struct pci_dev *pdev;
struct sg_table sgt;
void *buf[EP_SG_SEGMENTS];
};
static int ep_sgdma_build_and_map(struct ep_sgdma_priv *priv)
{
struct scatterlist *sg;
int i, ret;
ret = sg_alloc_table(&priv->sgt, EP_SG_SEGMENTS, GFP_KERNEL);
if (ret)
return ret;
for_each_sg(priv->sgt.sgl, sg, EP_SG_SEGMENTS, i) {
priv->buf[i] = kmalloc(EP_SEG_SIZE, GFP_KERNEL | GFP_DMA);
if (!priv->buf[i]) {
ret = -ENOMEM;
goto err_free_bufs;
}
memset(priv->buf[i], 0xA0 + i, EP_SEG_SIZE);
sg_set_buf(sg, priv->buf[i], EP_SEG_SIZE);
}
ret = dma_map_sgtable(&priv->pdev->dev, &priv->sgt, DMA_TO_DEVICE, 0);
if (ret) {
dev_err(&priv->pdev->dev, "dma_map_sgtable failed: %d\n", ret);
goto err_free_bufs;
}
for_each_sgtable_dma_sg(&priv->sgt, sg, i) {
dma_addr_t addr = sg_dma_address(sg);
unsigned int len = sg_dma_len(sg);
dev_info(&priv->pdev->dev,
"ep_sgdma: segment %d -> dma_addr=%pad len=%u\n",
i, &addr, len);
}
return 0;
err_free_bufs:
while (--i >= 0)
kfree(priv->buf[i]);
sg_free_table(&priv->sgt);
return ret;
}
static void ep_sgdma_unmap_and_free(struct ep_sgdma_priv *priv)
{
int i;
dma_unmap_sgtable(&priv->pdev->dev, &priv->sgt, DMA_TO_DEVICE, 0);
sg_free_table(&priv->sgt);
for (i = 0; i buf[i]);
}
static int ep_pci_sgdma_probe(struct pci_dev *pdev,
const struct pci_device_id *id)
{
struct ep_sgdma_priv *priv;
int ret;
priv = devm_kzalloc(&pdev->dev, sizeof(*priv), GFP_KERNEL);
if (!priv)
return -ENOMEM;
priv->pdev = pdev;
ret = pcim_enable_device(pdev);
if (ret)
return ret;
pci_set_master(pdev);
ret = dma_set_mask_and_coherent(&pdev->dev, DMA_BIT_MASK(28));
if (ret) {
dev_err(&pdev->dev, "no usable DMA mask\n");
return ret;
}
pci_set_drvdata(pdev, priv);
ret = ep_sgdma_build_and_map(priv);
if (ret)
return ret;
dev_info(&pdev->dev, "ep_pci_sgdma_demo: probe complete, %u segments mapped\n",
priv->sgt.nents);
return 0;
}
static void ep_pci_sgdma_remove(struct pci_dev *pdev)
{
struct ep_sgdma_priv *priv = pci_get_drvdata(pdev);
ep_sgdma_unmap_and_free(priv);
dev_info(&pdev->dev, "ep_pci_sgdma_demo: unmapped and cleaned up\n");
}
static const struct pci_device_id ep_sgdma_ids[] = {
{ PCI_DEVICE(0x1234, 0x11e8) }, /* QEMU edu device */
{ }
};
MODULE_DEVICE_TABLE(pci, ep_sgdma_ids);
static struct pci_driver ep_pci_sgdma_driver = {
.name = "ep_pci_sgdma_demo",
.id_table = ep_sgdma_ids,
.probe = ep_pci_sgdma_probe,
.remove = ep_pci_sgdma_remove,
};
module_pci_driver(ep_pci_sgdma_driver);
MODULE_LICENSE("GPL");
MODULE_DESCRIPTION("EmbeddedPathashala scatter-gather DMA demo for the free Linux kernel development course");
Build and Run Walkthrough
Build the module against your running kernel’s headers, exactly as in earlier lectures of this free linux development course:
make -C /lib/modules/$(uname -r)/build M=$(pwd) modules
Boot QEMU with the edu device attached (same flag used across this series), then load the module:
qemu-system-x86_64 -device edu -kernel bzImage -append "console=ttyS0" ...
sudo insmod ep_pci_sgdma_demo.ko
dmesg | tail -n 6
Expected output on probe — three distinct device-visible bus addresses, one per segment:
[ 12.104332] ep_pci_sgdma_demo 0000:00:03.0: ep_sgdma: segment 0 -> dma_addr=0x00000000ff01000 len=4096
[ 12.104351] ep_pci_sgdma_demo 0000:00:03.0: ep_sgdma: segment 1 -> dma_addr=0x00000000ff02000 len=4096
[ 12.104367] ep_pci_sgdma_demo 0000:00:03.0: ep_sgdma: segment 2 -> dma_addr=0x00000000ff03000 len=4096
[ 12.104380] ep_pci_sgdma_demo 0000:00:03.0: ep_pci_sgdma_demo: probe complete, 3 segments mapped
On rmmod, the driver unmaps and frees cleanly:
sudo rmmod ep_pci_sgdma_demo
dmesg | tail -n 2
[ 40.552110] ep_pci_sgdma_demo 0000:00:03.0: ep_pci_sgdma_demo: unmapped and cleaned up
Common Mistakes and Troubleshooting
- Unmapping with orig_nents instead of nents: if an IOMMU merged segments during mapping, unmapping with the pre-map count corrupts IOMMU state.
dma_unmap_sgtable()avoids this by readingnentsfrom the table itself. - Writing to
sg->page_linkordma_addressdirectly: always go throughsg_set_page()/sg_set_buf()andsg_dma_address()/sg_dma_len()— the internal encoding is not a plain pointer. - Forgetting
sg_free_table(): unmapping releases the DMA/IOMMU mapping but not the scatterlist array itself; both calls are required. - Touching CPU buffers after mapping without a sync call: once mapped for
DMA_TO_DEVICE/DMA_FROM_DEVICE, the buffer is owned by the device until you sync it back, exactly as with single-buffer streaming mappings. - Ignoring the return value of
dma_map_sgtable(): a mapping failure (for example, IOMMU exhaustion) must be handled — never assume the call always succeeds.
Best Practices
- Prefer
dma_map_sgtable()/dma_unmap_sgtable()over the rawdma_map_sg()/dma_unmap_sg()pair in all new driver code — the table-based API keeps entry counts consistent automatically. - Always iterate the mapped list with
for_each_sgtable_dma_sg(), never a hand-rolled loop over the raw array, so IOMMU-merged entries are handled correctly. - Keep the number of segments as small as your hardware descriptor ring realistically allows — every segment still costs one descriptor slot even though the mapping call itself is batched.
- For performance-sensitive paths, avoid rebuilding the scatter-gather table on every transfer; reuse it and rely on the sync calls where the buffer set doesn’t change between transfers.
- For security, never trust a scatter-gather length or offset that originated from user space without validating it against the actual allocated buffer size before calling
sg_set_buf()orsg_set_page().
Summary and Key Takeaways
Scatter-gather DMA lets a driver describe many physically scattered buffers as one logical transfer instead of issuing one single-buffer mapping per segment. struct sg_table tracks both the original segment count and the post-mapping count, and the modern dma_map_sgtable() / dma_unmap_sgtable() pair keeps those two numbers consistent automatically, which is why they’ve replaced the older raw dma_map_sg() call in current kernel code. Combined with for_each_sgtable_dma_sg(), sg_dma_address(), and sg_dma_len(), this gives a portable, IOMMU-safe way to hand a device a full descriptor list in one shot. This closes out the DMA portion of the PCI driver series in this free linux kernel development course — the next lecture moves on to the NVMEM framework.
Frequently Asked Questions
What is the difference between dma_map_sg() and dma_map_sgtable()?
dma_map_sg() returns the mapped entry count as a plain integer that the driver must store back manually. dma_map_sgtable() operates on the whole struct sg_table and updates its nents field internally, keeping the original and mapped counts self-consistent — this is why it’s the recommended API for new drivers.
Why can the mapped entry count be lower than the number of buffers I allocated?
On platforms with an IOMMU, physically adjacent segments can be merged into a single device-visible bus address during mapping, which reduces the effective entry count. This is exactly why iteration must use sgt.nents, not the original allocation count.
Can I write directly to a scatterlist entry’s dma_address field?
No. Always use the sg_dma_address() and sg_dma_len() accessor macros. The exact storage location of these values inside the structure is architecture-dependent, so reading or writing the field directly breaks portability.
Do I still need sg_free_table() after calling dma_unmap_sgtable()?
Yes. Unmapping releases the DMA/IOMMU mapping only. sg_free_table() releases the scatterlist array itself (and any chained pages it required), and both calls are needed for a clean teardown.
When should I use sg_set_buf() versus sg_set_page()?
Use sg_set_buf() when you already have a kernel virtual address, such as from kmalloc() — it derives the page and offset for you. Use sg_set_page() when you already hold struct page * pointers directly, which is common when working with pinned user pages.
Is scatter-gather DMA only useful for large transfers?
It’s most valuable whenever data spans multiple physically non-contiguous buffers, regardless of total size — network packet fragments, multi-segment block I/O requests, and video frame buffers all commonly use it even for moderate transfer sizes.
What happens if dma_map_sgtable() fails partway through mapping?
It returns a non-zero error code and does not leave a partially-mapped table behind — treat any non-zero return as a full failure, free the scatterlist table, and release the buffers you allocated before returning the error up the probe path.
Does this lecture belong to a free course?
Yes. This lecture is part of EmbeddedPathashala’s free Linux kernel development course, which also covers character devices, platform drivers, DMA, interrupts, and PCI/PCIe as part of a broader free embedded systems course and free Linux device drivers course.
Continue the Free Linux Kernel Development Course
Next up: the NVMEM framework for non-volatile storage devices like EEPROM.
Next Lecture Browse Full Course Index