V4L2 Buffer Dequeuing Explained-Free Linux Device Drivers Course

V4L2 Buffer Dequeuing Explained | EmbeddedPathashala
Free Linux Kernel Development Course

V4L2 Buffer Dequeuing Explained

A hands-on lecture from EmbeddedPathashala’s free Linux device drivers course on how V4L2 buffer dequeuing works in user space, from MMAP and USERPTR to modern DMABUF, with a complete working capture program.

Topics Covered In This Lecture

V4L2 buffer dequeuing VIDIOC_QBUF / VIDIOC_DQBUF MMAP streaming I/O USERPTR buffers DMABUF zero-copy videobuf2 framework

What Is V4L2 Buffer Dequeuing?

V4L2 buffer dequeuing is the mechanism by which a user-space application retrieves a filled video frame from the kernel’s Video4Linux2 (V4L2) subsystem after a capture device — a USB webcam, a CSI camera sensor, or an HDMI capture card — has written pixel data into a shared buffer. Every V4L2 capture driver in the Linux kernel, whether it is uvcvideo for USB cameras or a platform-specific sensor driver built on the videobuf2 (vb2) framework, exposes the same streaming I/O contract to user space: you queue empty buffers with VIDIOC_QBUF, the driver fills them as frames arrive, and you dequeue the completed buffer with VIDIOC_DQBUF.

This lecture is part of EmbeddedPathashala’s free Linux device drivers course, and it walks through the three buffer-passing strategies the V4L2 API supports — memory-mapped (MMAP) buffers, user-allocated pointer (USERPTR) buffers, and the modern zero-copy DMABUF method — along with the legacy read()/write() path. By the end, you will have built and run a complete, original capture program against a real /dev/videoX device.

What You Will Learn

This lecture from our free embedded Linux course covers:

V4L2 buffer lifecycle VIDIOC_QBUF / VIDIOC_DQBUF MMAP capture loop USERPTR buffer matching DMABUF zero-copy poll() based readiness videobuf2 kernel internals Writing a capture demo

Prerequisites

Before this lecture, you should be comfortable with C on Linux, basic file descriptor operations (open(), ioctl(), close()), and ideally have already covered the earlier V4L2 lectures in this series on device enumeration and format negotiation. A Linux machine with a USB webcam (or any UVC-compatible camera) and the v4l-utils package installed is recommended so you can follow along.

The V4L2 Streaming I/O Model

Before diving into buffer dequeuing specifically, it helps to see the full lifecycle a streaming V4L2 application follows. Every method — MMAP, USERPTR, or DMABUF — shares the same sequence of ioctl calls; only how the buffer memory itself is provided differs.

V4L2 Streaming Capture Lifecycle
1. open(“/dev/video0”) — obtain a file descriptor for the capture device
2. VIDIOC_REQBUFS — ask the driver to allocate N buffers (vb2 core in kernel)
3. VIDIOC_QUERYBUF / mmap() — map each buffer into user space (MMAP mode only)
4. VIDIOC_QBUF — queue every buffer so the driver can start filling it
5. VIDIOC_STREAMON — start the capture stream on the device
6. poll() — wait until a buffer is ready for reading
7. VIDIOC_DQBUF — dequeue the completed buffer and process the frame
8. VIDIOC_QBUF — requeue the same buffer, loop back to step 6
9. VIDIOC_STREAMOFF / close() — stop streaming and release resources

On the kernel side, this ioctl sequence is handled by the videobuf2 (vb2) core (drivers/media/common/videobuf2/), which every modern V4L2 capture driver builds on. The driver author only implements a small set of vb2 queue operations (queue_setup, buf_prepare, buf_queue, start_streaming, stop_streaming); vb2 itself handles the buffer state machine, DMA mapping, and the dequeue wait-queue that VIDIOC_DQBUF blocks on. Understanding the user-space dequeue call therefore also tells you what the kernel driver is doing underneath.

Dequeuing MMAP Buffers

MMAP is the default and most widely used streaming method. The kernel allocates the buffer memory (via VIDIOC_REQBUFS with memory = V4L2_MEMORY_MMAP), and user space maps each buffer into its own address space with mmap() using the offset returned by VIDIOC_QUERYBUF. Dequeuing then looks like this:

struct v4l2_buffer buf = {0};
buf.type   = V4L2_BUF_TYPE_VIDEO_CAPTURE;
buf.memory = V4L2_MEMORY_MMAP;

if (ioctl(fd, VIDIOC_DQBUF, &buf) == -1) {
    if (errno == EAGAIN)
        return 0;               /* no frame ready yet, try again later */
    perror("VIDIOC_DQBUF");
    exit(EXIT_FAILURE);
}

/* buf.index tells us which mapped buffer holds the frame */
ep_process_frame(mapped_buffers[buf.index].start, buf.bytesused);

/* hand the buffer back to the driver for the next frame */
if (ioctl(fd, VIDIOC_QBUF, &buf) == -1)
    perror("VIDIOC_QBUF");

The kernel guarantees buf.index always refers to a buffer you previously queued, so there is no manual address matching required — this is one reason MMAP is the easiest method to get right.

Dequeuing USERPTR Buffers

With V4L2_MEMORY_USERPTR, the application supplies its own heap-allocated buffers (for example via malloc() or posix_memalign()) instead of mapping kernel memory. This is useful when you want capture frames to land directly inside a buffer that another part of your pipeline (a neural network inference buffer, for instance) already owns. Because the driver only sees an address and a length, dequeuing requires you to match the returned pointer back to your own buffer table:

struct v4l2_buffer buf = {0};
buf.type   = V4L2_BUF_TYPE_VIDEO_CAPTURE;
buf.memory = V4L2_MEMORY_USERPTR;

if (ioctl(fd, VIDIOC_DQBUF, &buf) == -1) {
    if (errno == EAGAIN) return 0;
    perror("VIDIOC_DQBUF");
    exit(EXIT_FAILURE);
}

int idx = -1;
for (int i = 0; i < EP_BUF_COUNT; i++) {
    if (buf.m.userptr == (unsigned long)ep_bufs[i].start &&
        buf.length == ep_bufs[i].length) {
        idx = i;
        break;
    }
}
assert(idx != -1);
ep_process_frame((void *)buf.m.userptr, buf.bytesused);
ioctl(fd, VIDIOC_QBUF, &buf);

DMABUF On Modern Kernels

The book this lecture is modernized from predates widespread DMABUF adoption. On current kernels, V4L2_MEMORY_DMABUF is the preferred method whenever a frame needs to flow between two hardware-accelerated subsystems without a CPU copy — for example, handing a captured frame straight to a GPU shader, a hardware JPEG/H.264 encoder, or an NPU for on-device AI inference. Instead of mapping memory, you export or import a dma_buf file descriptor:

struct v4l2_buffer buf = {0};
buf.type   = V4L2_BUF_TYPE_VIDEO_CAPTURE;
buf.memory = V4L2_MEMORY_DMABUF;
buf.index  = i;
buf.m.fd   = dmabuf_fd[i];   /* fd from the exporting driver, e.g. a V4L2 M2M encoder */

ioctl(fd, VIDIOC_QBUF, &buf);
/* ... later ... */
ioctl(fd, VIDIOC_DQBUF, &buf); /* buf.m.fd still identifies the same dma_buf */

Because no CPU copy occurs, DMABUF pipelines are significantly more power- and bandwidth-efficient on embedded SoCs, which is why frameworks like GStreamer’s v4l2src/v4l2h264enc pair and most Linux camera stacks (libcamera included) default to it when the hardware supports buffer sharing.

Read/Write I/O Mode

The simplest — and slowest — method skips buffer queues entirely and uses a plain read() system call:

ssize_t n = read(fd, frame_buf, frame_buf_len);
if (n == -1) {
    if (errno == EAGAIN) return 0;
    perror("read");
    exit(EXIT_FAILURE);
}
ep_process_frame(frame_buf, (size_t)n);

Read/write mode is optional in the V4L2 spec, many modern drivers (including uvcvideo) do not implement it at all, and it always involves a kernel-to-user copy. It is only worth reaching for in throwaway scripts or devices that genuinely lack streaming support.

Comparison Table: Choosing A Buffer Method

MethodBuffer OwnerZero-CopyTypical Use Case
MMAPKernel (vb2)Yes (CPU-mapped)General-purpose capture apps, default choice
USERPTRUser spaceYes, if page-alignedFeeding an existing app-owned buffer
DMABUFAnother driver/GPUYes, true zero-copyCamera → encoder/NPU pipelines
read()/write()User spaceNoLegacy or unsupported by many drivers

Building A Demo Capture Program

Let’s put V4L2 buffer dequeuing into practice with an original, minimal MMAP capture program, ep_v4l2_grab.c. It opens a camera, queues four buffers, streams exactly ten frames, writes the last one to disk as a raw file, and cleans up.

/* ep_v4l2_grab.c — EmbeddedPathashala V4L2 MMAP capture demo */
#include <stdio.h>
#include <stdlib.h>
#include <string.h>
#include <errno.h>
#include <fcntl.h>
#include <unistd.h>
#include <sys/ioctl.h>
#include <sys/mman.h>
#include <linux/videodev2.h>

#define EP_DEVICE     "/dev/video0"
#define EP_BUF_COUNT  4
#define EP_FRAMES     10

struct ep_buffer { void *start; size_t length; };
static struct ep_buffer ep_bufs[EP_BUF_COUNT];

static void ep_die(const char *msg) { perror(msg); exit(EXIT_FAILURE); }

int main(void)
{
    int fd = open(EP_DEVICE, O_RDWR);
    if (fd == -1) ep_die("open");

    struct v4l2_format fmt = {0};
    fmt.type = V4L2_BUF_TYPE_VIDEO_CAPTURE;
    fmt.fmt.pix.width       = 640;
    fmt.fmt.pix.height      = 480;
    fmt.fmt.pix.pixelformat = V4L2_PIX_FMT_YUYV;
    fmt.fmt.pix.field       = V4L2_FIELD_NONE;
    if (ioctl(fd, VIDIOC_S_FMT, &fmt) == -1) ep_die("VIDIOC_S_FMT");

    struct v4l2_requestbuffers req = {0};
    req.count  = EP_BUF_COUNT;
    req.type   = V4L2_BUF_TYPE_VIDEO_CAPTURE;
    req.memory = V4L2_MEMORY_MMAP;
    if (ioctl(fd, VIDIOC_REQBUFS, &req) == -1) ep_die("VIDIOC_REQBUFS");

    for (int i = 0; i < EP_BUF_COUNT; i++) {
        struct v4l2_buffer buf = {0};
        buf.type   = V4L2_BUF_TYPE_VIDEO_CAPTURE;
        buf.memory = V4L2_MEMORY_MMAP;
        buf.index  = i;
        if (ioctl(fd, VIDIOC_QUERYBUF, &buf) == -1) ep_die("VIDIOC_QUERYBUF");

        ep_bufs[i].length = buf.length;
        ep_bufs[i].start  = mmap(NULL, buf.length, PROT_READ | PROT_WRITE,
                                  MAP_SHARED, fd, buf.m.offset);
        if (ep_bufs[i].start == MAP_FAILED) ep_die("mmap");
        if (ioctl(fd, VIDIOC_QBUF, &buf) == -1) ep_die("VIDIOC_QBUF");
    }

    enum v4l2_buf_type type = V4L2_BUF_TYPE_VIDEO_CAPTURE;
    if (ioctl(fd, VIDIOC_STREAMON, &type) == -1) ep_die("VIDIOC_STREAMON");

    for (int frame = 0; frame < EP_FRAMES; frame++) {
        struct v4l2_buffer buf = {0};
        buf.type   = V4L2_BUF_TYPE_VIDEO_CAPTURE;
        buf.memory = V4L2_MEMORY_MMAP;

        if (ioctl(fd, VIDIOC_DQBUF, &buf) == -1) ep_die("VIDIOC_DQBUF");
        printf("captured frame %d: %u bytes (buffer index %u)\n",
               frame, buf.bytesused, buf.index);

        if (frame == EP_FRAMES - 1) {
            FILE *out = fopen("ep_frame_last.yuv", "wb");
            fwrite(ep_bufs[buf.index].start, 1, buf.bytesused, out);
            fclose(out);
        }

        if (ioctl(fd, VIDIOC_QBUF, &buf) == -1) ep_die("VIDIOC_QBUF");
    }

    ioctl(fd, VIDIOC_STREAMOFF, &type);
    for (int i = 0; i < EP_BUF_COUNT; i++)
        munmap(ep_bufs[i].start, ep_bufs[i].length);
    close(fd);
    return 0;
}

Build it:

$ gcc -Wall -O2 -o ep_v4l2_grab ep_v4l2_grab.c
$ ./ep_v4l2_grab

Expected output:

captured frame 0: 614400 bytes (buffer index 1)
captured frame 1: 614400 bytes (buffer index 2)
captured frame 2: 614400 bytes (buffer index 3)
captured frame 3: 614400 bytes (buffer index 0)
captured frame 4: 614400 bytes (buffer index 1)
captured frame 5: 614400 bytes (buffer index 2)
captured frame 6: 614400 bytes (buffer index 3)
captured frame 7: 614400 bytes (buffer index 0)
captured frame 8: 614400 bytes (buffer index 1)
captured frame 9: 614400 bytes (buffer index 2)
$ ls -la ep_frame_last.yuv
-rw-r--r-- 1 user user 614400 ep_frame_last.yuv

614,400 bytes matches 640 × 480 × 2 bytes/pixel for YUYV — exactly what we expect for an uncompressed 4:2:2 frame at that resolution. If you enable kernel tracing (covered later in this series), you’ll see matching vb2_core_qbuf/vb2_core_dqbuf lines in dmesg for every frame.

Real-World Use Cases

V4L2 buffer dequeuing sits underneath almost every Linux camera application you have used. GStreamer’s v4l2src element performs exactly this MMAP queue/dequeue loop internally before handing frames to the rest of the pipeline. libcamera, the modern camera stack used on many ARM SoCs and Raspberry Pi boards, uses DMABUF buffer sharing between the ISP driver and downstream encoders for zero-copy video pipelines. Industrial machine-vision applications built on V4L2 (e.g. GenICam-to-V4L2 bridges) rely on USERPTR to capture frames directly into pre-allocated, page-locked memory for real-time inspection systems. Even OBS Studio’s Linux video-capture source is, at its core, this same dequeue loop.

Common Mistakes And Troubleshooting

Forgetting to check EAGAIN. If you open the device in non-blocking mode, VIDIOC_DQBUF returns -1 with errno == EAGAIN when no frame is ready yet — this is not an error, it means “call poll() or try again.”

Never calling VIDIOC_STREAMON. A very common beginner mistake is queuing buffers and then immediately calling VIDIOC_DQBUF without starting the stream — the call will simply block forever since the driver was never told to start capturing.

Mismatched buf.memory. The memory field in struct v4l2_buffer must match what you used in VIDIOC_REQBUFS on every subsequent call, or the ioctl fails with EINVAL.

Not requeuing the buffer. If you dequeue a buffer, process it, and forget the follow-up VIDIOC_QBUF, you will run out of free buffers after EP_BUF_COUNT frames and streaming will silently stall.

Best Practices

Use poll() (or select()) on the device file descriptor before calling VIDIOC_DQBUF rather than busy-looping — it lets your process sleep until a frame is actually ready and keeps CPU usage near zero between frames. Prefer MMAP unless you have a specific reason for USERPTR or DMABUF: it has the widest driver support and the fewest edge cases. Request at least three or four buffers (VIDIOC_REQBUFS count) so the driver always has a free buffer to fill while your application processes the previous one — two buffers is the bare minimum and leaves no slack for scheduling jitter.

Performance Considerations

Every MMAP/USERPTR frame still costs a page-table walk and, for USERPTR, a page-pinning operation on each queue/dequeue cycle; DMABUF avoids this entirely when the whole pipeline stays in kernel space. On resource-constrained SoCs, increasing buffer count trades memory for smoother frame delivery under scheduling pressure.

Security Considerations

Camera device nodes are typically restricted to the video group — never widen /dev/video* permissions globally. When using USERPTR, validate that buffer lengths you pass to the kernel match your actual allocation; passing an undersized buffer can lead to a driver writing past the end of user memory on older, less-hardened drivers.

Summary And Key Takeaways

V4L2 buffer dequeuing is the core mechanism every streaming capture application relies on: queue buffers, start the stream, dequeue completed frames in a loop, requeue, and stop. MMAP is the default and safest starting point, USERPTR lets you capture directly into application-owned memory, and DMABUF is the modern, zero-copy choice for camera-to-accelerator pipelines on today’s embedded Linux systems. The vb2 framework in the kernel implements the same state machine on the driver side, so understanding the user-space ioctl sequence directly maps to how a V4L2 driver behaves internally.

Conclusion

You’ve now built and run a real V4L2 MMAP capture program from scratch, and you understand how the three buffer-passing strategies differ and when to reach for each one. This is foundational knowledge for anyone working on Linux camera drivers, embedded vision pipelines, or multimedia frameworks. In the next lecture in this free Linux kernel development course, we move from writing capture code by hand to controlling and inspecting V4L2 devices interactively using the v4l2-ctl command-line tool.

Frequently Asked Questions

What does VIDIOC_DQBUF actually do in the Linux kernel?

It asks the V4L2 core (backed by the videobuf2 framework) to hand back the oldest completed buffer from the driver’s “done” queue, blocking (or returning EAGAIN in non-blocking mode) until one is available.

Why does my VIDIOC_DQBUF call block forever?

The most common cause is forgetting to call VIDIOC_STREAMON after queuing your initial buffers, or the device simply never producing frames (wrong input/format selected).

Is MMAP or DMABUF faster for V4L2 capture?

MMAP already avoids a full memory copy since it maps kernel buffer memory into user space, but DMABUF goes further by letting the buffer be shared directly with another driver (like a hardware encoder) with zero CPU involvement at all.

Do all V4L2 drivers support USERPTR and DMABUF?

No — support is optional per driver and per queue. Check the capabilities reported by VIDIOC_REQBUFS or VIDIOC_QUERYCAP; most modern vb2-based drivers support MMAP and DMABUF, while USERPTR support varies.

What is the difference between V4L2_MEMORY_MMAP and read()?

read() is a simple blocking system call with no buffer queue and always copies data into your buffer; MMAP uses an explicit queue/dequeue model that supports multiple in-flight buffers and non-blocking, poll()-driven capture.

How many buffers should I request with VIDIOC_REQBUFS?

Three or four is a reasonable default for most applications — enough to absorb scheduling jitter without wasting significant memory, though high frame-rate or high-latency-processing pipelines may need more.

Where can I learn more about the videobuf2 framework?

The official Linux kernel documentation at docs.kernel.org includes a dedicated Media subsystem section covering the videobuf2 core API used by nearly all modern V4L2 capture drivers.

Is this lecture part of a full Linux device driver course?

Yes — this is one lecture in EmbeddedPathashala’s free Linux kernel development course, which also covers character drivers, platform drivers, DMA, interrupts, and the V4L2 subsystem in depth.

Continue Learning Linux Kernel Development — Free

This lecture is part of EmbeddedPathashala’s free Linux kernel development course and free Linux device drivers course, built for engineers who want real embedded Linux skills without paying for a bootcamp.

Browse The Full Course Join The Community

Leave a Reply

Your email address will not be published. Required fields are marked *