V4L2 Buffer Queueing And Streaming-Free Linux Device Drivers Course

V4L2 Buffer Queueing And Streaming
Free Linux Kernel Development Course — V4L2 User Space API, Chapter 9 (Part 6 of 10)
VIDIOC_QBUF
VIDIOC_STREAMON
DMABUF / PRIME
free linux kernel development course
free linux device drivers course
free embedded linux course
v4l2 buffer queueing
VIDIOC_QBUF VIDIOC_STREAMON

In this lecture of our free Linux kernel development course, we tackle one of the most misunderstood steps in V4L2 video capture programming: V4L2 buffer queueing with VIDIOC_QBUF, and turning the pipeline on with VIDIOC_STREAMON. Once you understand how a buffer moves from “empty and idle” to “locked, DMA-mapped, and owned by the driver,” the rest of the V4L2 capture loop becomes trivial. This lecture is part of EmbeddedPathashala’s free embedded systems course and continues our hands-on free linux device drivers course series on the V4L2 subsystem.

What You Will Learn

  • What buffer enqueueing actually does at the kernel level, and why it locks memory
  • The exact semantics of VIDIOC_QBUF for MMAP, USERPTR, and DMABUF buffers
  • The “prime buffers” pattern used before entering a real capture loop
  • How DMABUF import works with the DRM PRIME framework on modern kernels
  • How VIDIOC_STREAMON flips a video queue into streaming state
  • Writing and running an original demo, ep_v4l2_qstream, on a real V4L2 node

Prerequisites

You should already be comfortable with VIDIOC_REQBUFS and VIDIOC_QUERYBUF /mmap(), covered earlier in this free linux kernel development course V4L2 series. If you haven’t allocated and mapped your buffers yet, go back to the previous lecture in this chapter before continuing here.

Why A Buffer Must Be Enqueued Before It Can Be Filled

A V4L2 buffer allocated with VIDIOC_REQBUFS starts life in an idle state — the kernel knows it exists, but it isn’t part of any active transfer. Before the driver’s hardware (usually a DMA engine behind a camera sensor, CSI bridge, or capture card) can write frame data into that buffer, the application must hand ownership of the buffer over to the driver. That handover is exactly what VIDIOC_QBUF does — it places the buffer on the driver’s internal queue and marks it as “in flight.”

This handover has a real consequence: the kernel pins the buffer’s physical pages in RAM so they cannot be swapped out to disk while DMA is in progress. A DMA engine works with physical addresses, not virtual ones, and it has no idea what a page fault is — if the page it’s writing to were paged out mid-transfer, you’d get silent memory corruption. This is why every buffer you enqueue stays locked until it is dequeued with VIDIOC_DQBUF, until VIDIOC_STREAMOFF or a fresh VIDIOC_REQBUFS is issued, or until the device node is closed.

One important rule that trips up a lot of beginners in this free linux device drivers course: once a buffer is queued and locked, your application must not read or write that buffer’s memory. The result of touching a locked buffer is undefined — the driver may be mid-DMA into it at any point.

Buffer Lifecycle State Flow
IDLE (after REQBUFS) → QUEUED (VIDIOC_QBUF) → LOCKED / DMA TARGET → FILLED → DEQUEUED (VIDIOC_DQBUF)

The Prime Buffers Pattern

For any capture application, it’s standard practice to enqueue every allocated buffer — commonly referred to as “priming” the queue — immediately after allocation, before you enable streaming and enter your read loop. If you only enqueue one buffer at a time and wait for it to fill before queueing the next, you introduce gaps where the hardware has nowhere to write, which shows up as dropped frames or stutter. Priming the full pool up front gives the driver a deep enough queue to stay busy continuously.

struct v4l2_buffer Fields Required For Enqueueing

Regardless of memory type, three fields are always mandatory when you call VIDIOC_QBUF: type (the buffer type, e.g. V4L2_BUF_TYPE_VIDEO_CAPTURE), memory (which I/O method this buffer uses), and index (which buffer slot from the pool you’re enqueueing). Everything else depends on the memory type you chose at VIDIOC_REQBUFS time.

Enqueueing MMAP Buffers On A Modern Kernel

MMAP is the most common path for capture applications, since the kernel owns the buffer memory and simply exposes it to user space via mmap(). Enqueueing one is the simplest case — no address or file descriptor needs to be supplied, since the kernel already knows where the memory lives.

#include <linux/videodev2.h>
#include <sys/ioctl.h>
#include <errno.h>

static int ep_v4l2_qbuf_mmap(int fd, unsigned int index)
{
    struct v4l2_buffer buf;

    memset(&buf, 0, sizeof(buf));
    buf.type   = V4L2_BUF_TYPE_VIDEO_CAPTURE;
    buf.memory = V4L2_MEMORY_MMAP;
    buf.index  = index;

    if (ioctl(fd, VIDIOC_QBUF, &buf) == -1) {
        perror("VIDIOC_QBUF (MMAP)");
        return -errno;
    }
    return 0;
}

On kernels supporting the multi-planar API (used by most modern SoC ISPs), you additionally attach a struct v4l2_plane array via buf.m.planes and set buf.length to the plane count — the kernel fills each plane’s offset internally for MMAP, so no extra work is needed beyond wiring up the array.

Enqueueing USERPTR Buffers

With USERPTR, your application owns the memory (typically allocated with a page-aligned posix_memalign()), so you must tell the kernel exactly where that memory lives and how big it is on every single enqueue call.

static int ep_v4l2_qbuf_userptr(int fd, unsigned int index,
                                 void *addr, size_t len)
{
    struct v4l2_buffer buf;

    memset(&buf, 0, sizeof(buf));
    buf.type       = V4L2_BUF_TYPE_VIDEO_CAPTURE;
    buf.memory     = V4L2_MEMORY_USERPTR;
    buf.index      = index;
    buf.m.userptr  = (unsigned long)addr;
    buf.length     = len;

    if (ioctl(fd, VIDIOC_QBUF, &buf) == -1) {
        perror("VIDIOC_QBUF (USERPTR)");
        return -errno;
    }
    return 0;
}

Enqueueing DMABUF Buffers For Zero-Copy Pipelines

DMABUF is the memory type you reach for when you want zero-copy handoff between two devices — for example, an output device (a decoder, or a GPU render target) feeding a capture device directly, with no CPU memcpy in between. Instead of an address, you supply a file descriptor that refers to a dma-buf object shared across subsystems.

static int ep_v4l2_qbuf_dmabuf(int fd, unsigned int index, int dmabuf_fd)
{
    struct v4l2_buffer buf;

    memset(&buf, 0, sizeof(buf));
    buf.type   = V4L2_BUF_TYPE_VIDEO_CAPTURE;
    buf.memory = V4L2_MEMORY_DMABUF;
    buf.index  = index;
    buf.m.fd   = dmabuf_fd;

    if (ioctl(fd, VIDIOC_QBUF, &buf) == -1) {
        perror("VIDIOC_QBUF (DMABUF)");
        return -errno;
    }
    return 0;
}

Where does that dmabuf_fd come from if the source isn’t a V4L2 device at all — say, a DRM/KMS render target? This is where the DRM subsystem’s PRIME framework comes in. PRIME is the name of the dma-buf sharing layer built on top of GEM (DRM’s memory manager). The flow looks like this on current kernels: allocate a GEM buffer object through your DRM driver, export it to a dma-buf file descriptor with DRM_IOCTL_PRIME_HANDLE_TO_FD, then hand that fd straight to VIDIOC_QBUF as shown above. No pixel data is copied at any point — only the buffer’s ownership and synchronization move between subsystems.

DRM PRIME To V4L2 DMABUF Import
DRM GEM Buffer → PRIME_HANDLE_TO_FD → dma-buf FD → VIDIOC_QBUF (DMABUF)

Enabling Streaming With VIDIOC_STREAMON

Priming the queue only prepares buffers — it does not start the hardware. Streaming must be explicitly turned on with VIDIOC_STREAMON, which tells the videobuf2 core (the modern buffer-management layer behind every V4L2 driver) that this queue is now live. Internally, this is the point where the driver typically starts the underlying hardware pipeline — enabling the sensor, kicking off the ISP, or arming the DMA engine — and only after this call will VIDIOC_DQBUF ever return a filled buffer.

static int ep_v4l2_streamon(int fd)
{
    enum v4l2_buf_type type = V4L2_BUF_TYPE_VIDEO_CAPTURE;

    if (ioctl(fd, VIDIOC_STREAMON, &type) == -1) {
        perror("VIDIOC_STREAMON");
        return -errno;
    }
    return 0;
}

Note that VIDIOC_STREAMON and VIDIOC_STREAMOFF both take a pointer to a plain enum v4l2_buf_type, not a full struct v4l2_buffer — a common mixup for newcomers to this free linux kernel development course.

Memory Type Comparison For Enqueueing

Memory TypeBuffer OwnerFields Set On QBUFTypical Use Case
V4L2_MEMORY_MMAPKerneltype, memory, indexSimple capture apps
V4L2_MEMORY_USERPTRApplication+ m.userptr, lengthCustom allocators, pre-existing buffers
V4L2_MEMORY_DMABUFExporter device (GPU/decoder/etc.)+ m.fdZero-copy device-to-device pipelines

Original Demo: ep_v4l2_qstream

Here’s a small, self-contained demo that primes an MMAP-based buffer queue and enables streaming on any V4L2 capture node. Build it against a current kernel’s linux/videodev2.h header.

// ep_v4l2_qstream.c — prime buffers and start streaming
#include <stdio.h>
#include <string.h>
#include <errno.h>
#include <fcntl.h>
#include <unistd.h>
#include <sys/ioctl.h>
#include <linux/videodev2.h>

#define EP_BUF_COUNT 4

int main(void)
{
    int fd = open("/dev/video0", O_RDWR);
    if (fd < 0) { perror("open"); return 1; }

    /* Assumes VIDIOC_REQBUFS with count = EP_BUF_COUNT, memory = MMAP,
       has already been done earlier in your program. */

    for (unsigned int i = 0; i < EP_BUF_COUNT; i++) {
        struct v4l2_buffer buf;
        memset(&buf, 0, sizeof(buf));
        buf.type   = V4L2_BUF_TYPE_VIDEO_CAPTURE;
        buf.memory = V4L2_MEMORY_MMAP;
        buf.index  = i;

        if (ioctl(fd, VIDIOC_QBUF, &buf) == -1) {
            perror("VIDIOC_QBUF");
            close(fd);
            return 1;
        }
        printf("ep_v4l2_qstream: queued buffer index %u\n", i);
    }

    enum v4l2_buf_type type = V4L2_BUF_TYPE_VIDEO_CAPTURE;
    if (ioctl(fd, VIDIOC_STREAMON, &type) == -1) {
        perror("VIDIOC_STREAMON");
        close(fd);
        return 1;
    }
    printf("ep_v4l2_qstream: streaming enabled\n");

    close(fd);
    return 0;
}

Build and run:

$ gcc -o ep_v4l2_qstream ep_v4l2_qstream.c
$ ./ep_v4l2_qstream

Expected output on a working capture node with four buffers requested:

ep_v4l2_qstream: queued buffer index 0
ep_v4l2_qstream: queued buffer index 1
ep_v4l2_qstream: queued buffer index 2
ep_v4l2_qstream: queued buffer index 3
ep_v4l2_qstream: streaming enabled

If any VIDIOC_QBUF call fails with EINVAL, double-check that buf.memory matches exactly what you passed to VIDIOC_REQBUFS, and that buf.index is within the buffer count you actually requested.

Common Mistakes And Troubleshooting

  • Calling VIDIOC_STREAMON before any buffer is queued: streaming will turn on, but your first VIDIOC_DQBUF will simply block forever, since there’s nothing for hardware to fill.
  • Mismatched buf.memory: if the value passed to VIDIOC_QBUF doesn’t match the memory type used at VIDIOC_REQBUFS, the ioctl fails with -EINVAL.
  • Touching buffer memory after enqueueing: reading or writing a queued MMAP or USERPTR buffer before it’s dequeued produces undefined results — the DMA engine may be actively writing to it.
  • Forgetting the multi-planar array: on multi-planar capture types, omitting the v4l2_plane array on buf.m.planes will fail the enqueue with -EINVAL.
  • Passing a v4l2_buffer to VIDIOC_STREAMON/STREAMOFF: these two ioctls expect a bare enum v4l2_buf_type, not a full buffer structure.

Best Practices

  • Always prime the entire buffer pool before calling VIDIOC_STREAMON for smooth, gap-free capture.
  • Wrap every ioctl call in an EINTR-retry helper, since signals can interrupt blocking ioctls on a busy system.
  • Check driver capability flags (from VIDIOC_QUERYCAP) before assuming DMABUF or multi-planar support is available on a given device.
  • Never assume enqueue order equals fill order across all drivers — always trust the index field returned by VIDIOC_DQBUF, covered in the next lecture.

Interview Questions

Why does VIDIOC_QBUF lock a buffer’s memory pages?

Because the underlying transfer is DMA-based and DMA engines operate on physical addresses. If the OS swapped a page out mid-transfer, the DMA engine would keep writing to a physical address that no longer holds the intended buffer, silently corrupting memory. Locking (pinning) the pages guarantees the physical address stays valid for the whole transfer.

What’s the practical difference between USERPTR and DMABUF memory types?

USERPTR buffers are plain user-space memory the application allocates itself and hands to the kernel by address on every enqueue — there’s still a mapping step for the DMA engine internally. DMABUF buffers are pre-existing kernel objects shared by file descriptor across subsystems (like DRM/GEM), enabling true zero-copy sharing between two hardware devices without the CPU ever touching the pixel data.

What happens internally when VIDIOC_STREAMON is called?

The videobuf2 core transitions the queue’s state machine into “streaming,” and the driver’s start_streaming callback runs, which typically programs and enables the underlying hardware — sensor, ISP, or DMA channel — so that already-queued buffers begin actually being filled.

Summary And Key Takeaways

Buffer queueing is the handoff point where your application gives up temporary ownership of a buffer so the driver’s hardware can fill it via DMA. VIDIOC_QBUF locks that buffer’s memory and its exact fields depend on the memory type — MMAP needs almost nothing extra, USERPTR needs an address and length, and DMABUF needs a shared file descriptor, often produced by the DRM PRIME framework. Priming the whole buffer pool before calling VIDIOC_STREAMON keeps capture smooth, and streaming itself is a simple on/off switch on the queue’s state machine. With buffers queued and streaming enabled, the next step in this free linux kernel development course is dequeueing filled buffers — which is exactly where the next lecture picks up.

Frequently Asked Questions

What does VIDIOC_QBUF actually do in V4L2?

It enqueues a previously allocated buffer onto the driver’s active queue, locking its memory pages so hardware DMA can safely fill it.

Do I need to enqueue buffers when using read/write I/O?

No. The read/write I/O method bypasses the queueing mechanism entirely — enqueueing only applies to the streaming I/O methods: MMAP, USERPTR, and DMABUF.

What error does VIDIOC_DQBUF return if I forgot to queue a buffer first?

It returns -EINVAL, since there’s nothing on the queue to dequeue.

Can I call VIDIOC_STREAMON more than once without STREAMOFF?

No — calling it again on an already-streaming queue is a driver-dependent error condition and should be avoided; always pair STREAMON with a matching STREAMOFF before restarting.

Why prime the buffer queue before starting the capture loop?

Priming gives the hardware a deep enough backlog of empty buffers to write into continuously, preventing frame drops caused by the driver having nowhere to DMA new frame data.

Is DMABUF queueing available on every V4L2 driver?

No — check the device’s capability flags from VIDIOC_QUERYCAP and, ideally, attempt a small test allocation, since DMABUF support depends on both the exporting and importing drivers.

What is PRIME in the context of DRM and V4L2?

PRIME is the DRM subsystem’s dma-buf sharing layer, built on top of GEM, that lets a GPU or display buffer be exported as a file descriptor and imported directly into a V4L2 device using V4L2_MEMORY_DMABUF.

Continue The Free Linux Kernel Development Course

Next up: dequeueing filled buffers with VIDIOC_DQBUF and building the full capture read loop.

Next Lecture Back To Course Index

Leave a Reply

Your email address will not be published. Required fields are marked *