V4L2 MMAP and DMABUF Buffers-Free Linux Device Drivers Course

V4L2 MMAP and DMABUF Buffers

Free Linux Kernel Development Course — V4L2 Buffer Memory Mapping Explained

Lecture 9.5
V4L2 Buffer Management
~20 min read

Welcome back to our free Linux kernel development course. In the previous lecture of this free embedded Linux course, we allocated buffers with VIDIOC_REQBUFS and compared the three I/O memory models that V4L2 supports. In this lecture we go one level deeper into V4L2 buffer memory mapping: how an MMAP buffer is actually turned into a pointer your application can read pixels from, how DMABUF buffers are exported and shared between two V4L2 devices without a single memcpy, and how the much simpler read/write model allocates memory entirely in user space. This is core material for anyone building camera, capture, or media pipeline drivers as part of our free linux device drivers course.

Topics covered in this lecture

VIDIOC_QUERYBUF mmap() buffer mapping VIDIOC_EXPBUF DMABUF exporter/importer Read/Write I/O buffers V4L2 memory model comparison

What You Will Learn

  • How VIDIOC_QUERYBUF turns a requested buffer into a physical offset
  • How mmap() maps that offset into your process address space
  • What DMABUF exporter/importer really means between two devices
  • How VIDIOC_EXPBUF hands out a DMABUF file descriptor
  • Why read/write I/O needs no ioctl at all for buffer setup
  • A full working demo you can build and run on a real V4L2 node

Prerequisites

This lecture assumes you have already gone through VIDIOC_REQBUFS and the memory model comparison from the previous lecture in this free linux kernel development course. You should be comfortable with basic V4L2 concepts (capture devices, buffer queues, ioctl-based device control) and standard POSIX mmap(). A Linux machine with a V4L2 capture device (a UVC webcam works fine) and kernel headers installed is enough to follow along and build the demo yourself.

Recap: Where MMAP Buffers Stand After REQBUFS

When your application calls VIDIOC_REQBUFS with memory = V4L2_MEMORY_MMAP, the driver allocates a set of physically contiguous (or scatter-gather, depending on the DMA engine) buffers inside kernel space. At that point your application knows nothing except how many buffers it got back in req.count. It has no pointer, no address, nothing it can read pixel data from yet. Two more steps are needed before any capture loop can begin: discovering where each buffer lives (VIDIOC_QUERYBUF) and mapping that location into user space (mmap()). This two-step handshake is exactly what the V4L2 core requires for every MMAP buffer, and getting it right is the single most common source of bugs in first-time V4L2 capture applications.

MMAP Buffer Mapping Flow

[User Space] [V4L2 Core / Driver] VIDIOC_REQBUFS ——————-> allocate N buffers in kernel (memory=MMAP) return req.count for each buffer index i: VIDIOC_QUERYBUF ——————> fill v4l2_buffer.m.offset (index = i) map kernel buffer pages <—————– into process VMA store returned pointer in buf_addr[i].start

VIDIOC_QUERYBUF: Discovering the Buffer Offset

VIDIOC_QUERYBUF is the ioctl that turns a plain buffer index into something you can actually map. You fill in type, memory, and index in a struct v4l2_buffer, and the kernel fills the rest — most importantly length (the buffer size in bytes) and, for the MMAP case, m.offset (a fake file offset that the mmap() call below uses to identify which kernel buffer to map). This offset is not a real memory address; it is a token the V4L2 core uses internally to route the mmap() request to the correct buffer.

struct v4l2_buffer v4l2_buf;
memset(&v4l2_buf, 0, sizeof(v4l2_buf));

v4l2_buf.type   = V4L2_BUF_TYPE_VIDEO_CAPTURE;
v4l2_buf.memory = V4L2_MEMORY_MMAP;
v4l2_buf.index  = buf_index;

if (ioctl(fd, VIDIOC_QUERYBUF, &v4l2_buf) == -1) {
    perror("VIDIOC_QUERYBUF");
    exit(EXIT_FAILURE);
}

/* v4l2_buf.length      -> size of this buffer in bytes   */
/* v4l2_buf.m.offset    -> token to pass into mmap()      */

Once you have length and m.offset, the actual mapping is a single standard POSIX call:

void *start = mmap(NULL,                    /* let kernel pick address   */
                    v4l2_buf.length,          /* map exactly this many bytes */
                    PROT_READ | PROT_WRITE,   /* app reads and writes buffer */
                    MAP_SHARED,               /* changes visible to driver  */
                    fd,                        /* the video device fd       */
                    v4l2_buf.m.offset);        /* offset from QUERYBUF      */

if (start == MAP_FAILED) {
    perror("mmap");
    exit(EXIT_FAILURE);
}

MAP_SHARED is not optional here. With MAP_PRIVATE the kernel would give you a copy-on-write mapping, so writes (and the driver’s own DMA fills) would never be seen consistently on both sides. MAP_SHARED guarantees that when the driver’s DMA engine writes a captured frame into the buffer, your process sees that exact memory without any extra copy — this zero-copy behaviour is the entire reason MMAP I/O exists.

DMABUF: The Exporter and Importer Model

MMAP and USERPTR both describe buffers that live inside one device’s driver. DMABUF exists to solve a different problem: sharing one buffer between two different devices without ever copying the pixel data between them. This matters enormously on modern SoCs where a camera sensor driver, an image signal processor (ISP), a hardware scaler, and a display or encoder may all need to touch the same frame in sequence. Copying a 4K frame at each stage would burn memory bandwidth and add latency; DMABUF lets every stage operate on the same physical memory.

The Linux DMA-BUF framework defines two roles:

  • Exporter — the driver that originally allocated the buffer and hands out a file descriptor referencing it
  • Importer — any other driver (or subsystem) that receives that file descriptor and attaches to the same underlying memory

In V4L2 terms: you first allocate buffers on the exporting device using VIDIOC_REQBUFS with memory = V4L2_MEMORY_MMAP (MMAP is the only memory type from which buffers can be exported). You then call VIDIOC_EXPBUF once per buffer index to get back a DMABUF file descriptor. That fd can now be handed to a completely different V4L2 node — for example an output/m2m device — by queuing it there with memory = V4L2_MEMORY_DMABUF and placing the fd in v4l2_buffer.m.fd.

struct v4l2_exportbuffer expbuf;
memset(&expbuf, 0, sizeof(expbuf));

expbuf.type  = V4L2_BUF_TYPE_VIDEO_CAPTURE;
expbuf.index = buf_index;
expbuf.flags = O_RDWR;   /* import/export permission on the returned fd */

if (ioctl(fd, VIDIOC_EXPBUF, &expbuf) == -1) {
    perror("VIDIOC_EXPBUF");
    exit(EXIT_FAILURE);
}

int dmabuf_fd = expbuf.fd;   /* pass this fd to the importer device */

DMABUF Exporter / Importer Sharing

[Capture Device A – Exporter] [Output/M2M Device B – Importer] VIDIOC_REQBUFS (MMAP) VIDIOC_REQBUFS (DMABUF) | VIDIOC_EXPBUF(index=i) | dmabuf_fd ——————-> v4l2_buffer.m.fd = dmabuf_fd VIDIOC_QBUF on device B Same physical memory, zero copy between A and B

One important detail that trips people up: exporting a buffer with VIDIOC_EXPBUF and then immediately mapping that same fd on the very same capture device is pointless — you already have direct mmap() access to those buffers on the exporting device, as shown in the previous section. DMABUF earns its keep specifically when the buffer needs to travel to a second device, which is the normal case on camera and media pipelines with an ISP, scaler, or encoder stage.

Read/Write I/O: The Simplest Buffer Model

Not every V4L2 device supports streaming I/O. Some simple capture devices only implement the classic read()/write() system calls, and even on devices that do support streaming, read/write mode is useful for quick tools and scripts where you don’t want to manage buffer queues at all.

The buffer handling here is trivial by comparison: there is no VIDIOC_REQBUFS, no VIDIOC_QUERYBUF, no ioctl of any kind. The kernel is not involved in buffer allocation — you simply allocate ordinary heap memory in user space with malloc() and pass that buffer straight into read(). The driver copies captured data into it behind the scenes, which means read/write I/O always costs one extra memory copy compared to MMAP or DMABUF — a real cost on high-resolution or high-frame-rate capture, but perfectly fine for low bandwidth devices or occasional captures.

size_t buffer_size = 640 * 480 * 2; /* example: YUYV 4:2:2 frame size */

unsigned char *frame_buf = malloc(buffer_size);
if (!frame_buf) {
    perror("malloc");
    exit(EXIT_FAILURE);
}

ssize_t n = read(fd, frame_buf, buffer_size);
if (n == -1) {
    perror("read");
    exit(EXIT_FAILURE);
}
/* frame_buf now holds n bytes of captured frame data */

V4L2 Memory Model Comparison

Aspect MMAP USERPTR DMABUF Read/Write
Buffer allocated by Kernel driver Application (page-aligned) Exporting driver Application (malloc)
Requires VIDIOC_REQBUFS Yes Yes Yes (on exporter) No
Extra memory copy None None None One copy per frame
Cross-device sharing Not directly Not directly Yes, by design Not applicable
Typical use case Standard camera capture apps Buffers from another allocator (GPU, encoder) Camera-ISP-encoder pipelines Quick scripts, simple devices

Hands-On: ep_v4l2_mmapbuf Demo

This original demo, ep_v4l2_mmapbuf, requests MMAP buffers, queries and maps each one, and prints the mapped address and size — the exact groundwork every streaming capture application needs before it can start queuing and dequeuing frames.

/* ep_v4l2_mmapbuf.c - EmbeddedPathashala V4L2 MMAP buffer mapping demo */
#include <stdio.h>
#include <stdlib.h>
#include <string.h>
#include <fcntl.h>
#include <unistd.h>
#include <sys/mman.h>
#include <sys/ioctl.h>
#include <linux/videodev2.h>

#define EP_BUF_COUNT 4

struct ep_buffer {
    void   *start;
    size_t  length;
};

int main(int argc, char **argv)
{
    const char *dev_name = (argc > 1) ? argv[1] : "/dev/video0";
    int fd = open(dev_name, O_RDWR);
    if (fd == -1) {
        perror("open");
        return EXIT_FAILURE;
    }

    struct v4l2_requestbuffers req;
    memset(&req, 0, sizeof(req));
    req.count  = EP_BUF_COUNT;
    req.type   = V4L2_BUF_TYPE_VIDEO_CAPTURE;
    req.memory = V4L2_MEMORY_MMAP;

    if (ioctl(fd, VIDIOC_REQBUFS, &req) == -1) {
        perror("VIDIOC_REQBUFS");
        close(fd);
        return EXIT_FAILURE;
    }

    if (req.count < 2) {
        fprintf(stderr, "Insufficient buffer memory on %s\n", dev_name);
        close(fd);
        return EXIT_FAILURE;
    }

    struct ep_buffer *buffers = calloc(req.count, sizeof(*buffers));
    if (!buffers) {
        perror("calloc");
        close(fd);
        return EXIT_FAILURE;
    }

    for (unsigned int i = 0; i < req.count; i++) {
        struct v4l2_buffer buf;
        memset(&buf, 0, sizeof(buf));
        buf.type   = V4L2_BUF_TYPE_VIDEO_CAPTURE;
        buf.memory = V4L2_MEMORY_MMAP;
        buf.index  = i;

        if (ioctl(fd, VIDIOC_QUERYBUF, &buf) == -1) {
            perror("VIDIOC_QUERYBUF");
            close(fd);
            return EXIT_FAILURE;
        }

        buffers[i].length = buf.length;
        buffers[i].start = mmap(NULL, buf.length,
                                 PROT_READ | PROT_WRITE, MAP_SHARED,
                                 fd, buf.m.offset);

        if (buffers[i].start == MAP_FAILED) {
            perror("mmap");
            close(fd);
            return EXIT_FAILURE;
        }

        printf("Buffer %u mapped at %p, size %zu bytes\n",
               i, buffers[i].start, buffers[i].length);
    }

    for (unsigned int i = 0; i < req.count; i++)
        munmap(buffers[i].start, buffers[i].length);

    free(buffers);
    close(fd);
    return EXIT_SUCCESS;
}

Build and run

$ gcc -Wall -o ep_v4l2_mmapbuf ep_v4l2_mmapbuf.c
$ ./ep_v4l2_mmapbuf /dev/video0

Expected output

Buffer 0 mapped at 0x7f2c1a3b4000, size 614400 bytes
Buffer 1 mapped at 0x7f2c1a24b000, size 614400 bytes
Buffer 2 mapped at 0x7f2c1a1e2000, size 614400 bytes
Buffer 3 mapped at 0x7f2c1a179000, size 614400 bytes

Exact addresses and buffer size will differ on your machine depending on the ASLR layout and the resolution/pixel format currently set on the device with VIDIOC_S_FMT from an earlier lecture.

Common Mistakes and Troubleshooting

  • Calling mmap() before VIDIOC_QUERYBUF — you need buf.length and buf.m.offset first
  • Using MAP_PRIVATE instead of MAP_SHARED — breaks zero-copy DMA visibility
  • Forgetting munmap() on every buffer before closing the device fd
  • Trying VIDIOC_EXPBUF with memory=USERPTR — export only works from MMAP buffers
  • Assuming DMABUF gives you a CPU pointer directly — you still map it separately if you need CPU access
  • Not checking req.count after REQBUFS — the driver can silently grant fewer buffers than requested

Best Practices

  • Always check the actual req.count returned, never assume you got what you asked for
  • Track buffer length alongside the pointer — munmap() needs both
  • Prefer DMABUF only when a buffer genuinely needs to cross device boundaries
  • Reserve read/write I/O for low-bandwidth or one-off capture tools, not production pipelines
  • Close every DMABUF file descriptor once the importer device is done with it

Interview Questions

Why does mmap() need VIDIOC_QUERYBUF before it can map a buffer?

mmap() needs a length and an offset to know what to map and how much. VIDIOC_QUERYBUF is the ioctl that fills in v4l2_buffer.length and v4l2_buffer.m.offset for a given buffer index — without those two values there is nothing valid to pass to mmap() at all.

What problem does DMABUF solve that MMAP cannot?

MMAP buffers are private to the device that allocated them and mapped into one process. DMABUF lets the same physical memory be shared between two different drivers or devices — for example a camera sensor and a scaler — without copying data between them, which MMAP alone cannot do.

Why does read/write I/O always cost one extra memory copy compared to MMAP?

With read/write I/O, the buffer lives in ordinary user-space heap memory allocated with malloc(). The driver captures data into its own internal buffer first, then the kernel copies it into your read() buffer — there’s no direct DMA-to-userspace path like MMAP/DMABUF provide.

Summary and Key Takeaways

In this lecture of our free linux kernel development course we completed the MMAP buffer mapping flow with VIDIOC_QUERYBUF and mmap(), learned how DMABUF’s exporter/importer model shares buffers across devices with zero copy using VIDIOC_EXPBUF, and saw why read/write I/O is the simplest but least efficient memory model available in V4L2. With all four memory models now covered, we’re ready to move on to actually enqueueing and dequeueing these buffers to run a real capture loop — the subject of the next lecture in this free linux device drivers course.

Frequently Asked Questions

What is the difference between VIDIOC_QUERYBUF and VIDIOC_REQBUFS?

VIDIOC_REQBUFS requests that a set of buffers be allocated in the driver and returns only a count. VIDIOC_QUERYBUF is called afterward, once per buffer index, to retrieve the details (length and mmap offset) of one specific buffer already allocated.

Can I use VIDIOC_EXPBUF on USERPTR buffers?

No. VIDIOC_EXPBUF only works on buffers allocated with V4L2_MEMORY_MMAP, since it exports the kernel-allocated memory behind those buffers as a DMABUF file descriptor.

Do I still need mmap() if I use DMABUF buffers?

Only if your application itself needs CPU access to the pixel data. If the DMABUF is only being passed between two hardware-accelerated drivers, no CPU mapping is required at all.

Which V4L2 memory model gives the best performance for camera capture?

MMAP is the standard choice for most single-device capture applications because it avoids extra copies and needs no external buffer allocator. DMABUF becomes the better choice once a second device needs to consume the same frame.

Why did VIDIOC_REQBUFS grant fewer buffers than I requested?

Drivers are allowed to cap the buffer count based on available memory or hardware limits. This is exactly why applications must check req.count after the call and fail gracefully if it falls below their minimum usable count.

Is read/write I/O deprecated in modern V4L2 drivers?

No, it is still a valid and supported capability, but many modern high-throughput drivers only implement streaming I/O (MMAP/USERPTR/DMABUF) since read/write cannot avoid the extra copy.

What flags should I pass to VIDIOC_EXPBUF?

Typically O_RDWR so the resulting file descriptor can be both read from and written to by the importing driver, unless your use case only ever needs read-only access.

Continue Your Free Linux Kernel Development Course

Master V4L2 buffer management step by step with EmbeddedPathashala’s free embedded systems course.

Next Lecture: Enqueueing Buffers Back to Course Index

Leave a Reply

Your email address will not be published. Required fields are marked *