In this lecture of our free Linux kernel development course, we tackle one of the
most misunderstood steps in V4L2 video capture programming: V4L2 buffer queueing
with VIDIOC_QBUF, and turning the pipeline on with VIDIOC_STREAMON. Once
you understand how a buffer moves from “empty and idle” to “locked, DMA-mapped, and owned by the
driver,” the rest of the V4L2 capture loop becomes trivial. This lecture is part of
EmbeddedPathashala’s free embedded systems course and continues our
hands-on free linux device drivers course series on the V4L2 subsystem.
What You Will Learn
- What buffer enqueueing actually does at the kernel level, and why it locks memory
- The exact semantics of
VIDIOC_QBUFfor MMAP, USERPTR, and DMABUF buffers - The “prime buffers” pattern used before entering a real capture loop
- How DMABUF import works with the DRM PRIME framework on modern kernels
- How
VIDIOC_STREAMONflips a video queue into streaming state - Writing and running an original demo,
ep_v4l2_qstream, on a real V4L2 node
Prerequisites
You should already be comfortable with VIDIOC_REQBUFS and VIDIOC_QUERYBUF
/mmap(), covered earlier in this free linux kernel development course
V4L2 series. If you haven’t allocated and mapped your buffers yet, go back to the previous lecture
in this chapter before continuing here.
Why A Buffer Must Be Enqueued Before It Can Be Filled
A V4L2 buffer allocated with VIDIOC_REQBUFS starts life in an idle state — the kernel
knows it exists, but it isn’t part of any active transfer. Before the driver’s hardware (usually a
DMA engine behind a camera sensor, CSI bridge, or capture card) can write frame data into that
buffer, the application must hand ownership of the buffer over to the driver. That handover is
exactly what VIDIOC_QBUF does — it places the buffer on the driver’s internal queue and
marks it as “in flight.”
This handover has a real consequence: the kernel pins the buffer’s physical pages in RAM so they
cannot be swapped out to disk while DMA is in progress. A DMA engine works with physical addresses,
not virtual ones, and it has no idea what a page fault is — if the page it’s writing to were paged
out mid-transfer, you’d get silent memory corruption. This is why every buffer you enqueue stays
locked until it is dequeued with VIDIOC_DQBUF, until VIDIOC_STREAMOFF or a
fresh VIDIOC_REQBUFS is issued, or until the device node is closed.
One important rule that trips up a lot of beginners in this free linux device drivers course: once a buffer is queued and locked, your application must not read or write that buffer’s memory. The result of touching a locked buffer is undefined — the driver may be mid-DMA into it at any point.
The Prime Buffers Pattern
For any capture application, it’s standard practice to enqueue every allocated buffer — commonly referred to as “priming” the queue — immediately after allocation, before you enable streaming and enter your read loop. If you only enqueue one buffer at a time and wait for it to fill before queueing the next, you introduce gaps where the hardware has nowhere to write, which shows up as dropped frames or stutter. Priming the full pool up front gives the driver a deep enough queue to stay busy continuously.
struct v4l2_buffer Fields Required For Enqueueing
Regardless of memory type, three fields are always mandatory when you call VIDIOC_QBUF:
type (the buffer type, e.g. V4L2_BUF_TYPE_VIDEO_CAPTURE), memory
(which I/O method this buffer uses), and index (which buffer slot from the pool you’re
enqueueing). Everything else depends on the memory type you chose at VIDIOC_REQBUFS time.
Enqueueing MMAP Buffers On A Modern Kernel
MMAP is the most common path for capture applications, since the kernel owns the buffer memory
and simply exposes it to user space via mmap(). Enqueueing one is the simplest case —
no address or file descriptor needs to be supplied, since the kernel already knows where the memory
lives.
#include <linux/videodev2.h>
#include <sys/ioctl.h>
#include <errno.h>
static int ep_v4l2_qbuf_mmap(int fd, unsigned int index)
{
struct v4l2_buffer buf;
memset(&buf, 0, sizeof(buf));
buf.type = V4L2_BUF_TYPE_VIDEO_CAPTURE;
buf.memory = V4L2_MEMORY_MMAP;
buf.index = index;
if (ioctl(fd, VIDIOC_QBUF, &buf) == -1) {
perror("VIDIOC_QBUF (MMAP)");
return -errno;
}
return 0;
}
On kernels supporting the multi-planar API (used by most modern SoC ISPs), you additionally
attach a struct v4l2_plane array via buf.m.planes and set
buf.length to the plane count — the kernel fills each plane’s offset internally for
MMAP, so no extra work is needed beyond wiring up the array.
Enqueueing USERPTR Buffers
With USERPTR, your application owns the memory (typically allocated with a page-aligned
posix_memalign()), so you must tell the kernel exactly where that memory lives and how
big it is on every single enqueue call.
static int ep_v4l2_qbuf_userptr(int fd, unsigned int index,
void *addr, size_t len)
{
struct v4l2_buffer buf;
memset(&buf, 0, sizeof(buf));
buf.type = V4L2_BUF_TYPE_VIDEO_CAPTURE;
buf.memory = V4L2_MEMORY_USERPTR;
buf.index = index;
buf.m.userptr = (unsigned long)addr;
buf.length = len;
if (ioctl(fd, VIDIOC_QBUF, &buf) == -1) {
perror("VIDIOC_QBUF (USERPTR)");
return -errno;
}
return 0;
}
Enqueueing DMABUF Buffers For Zero-Copy Pipelines
DMABUF is the memory type you reach for when you want zero-copy handoff between two devices — for example, an output device (a decoder, or a GPU render target) feeding a capture device directly, with no CPU memcpy in between. Instead of an address, you supply a file descriptor that refers to a dma-buf object shared across subsystems.
static int ep_v4l2_qbuf_dmabuf(int fd, unsigned int index, int dmabuf_fd)
{
struct v4l2_buffer buf;
memset(&buf, 0, sizeof(buf));
buf.type = V4L2_BUF_TYPE_VIDEO_CAPTURE;
buf.memory = V4L2_MEMORY_DMABUF;
buf.index = index;
buf.m.fd = dmabuf_fd;
if (ioctl(fd, VIDIOC_QBUF, &buf) == -1) {
perror("VIDIOC_QBUF (DMABUF)");
return -errno;
}
return 0;
}
Where does that dmabuf_fd come from if the source isn’t a V4L2 device at all — say,
a DRM/KMS render target? This is where the DRM subsystem’s PRIME framework comes
in. PRIME is the name of the dma-buf sharing layer built on top of GEM (DRM’s memory manager). The
flow looks like this on current kernels: allocate a GEM buffer object through your DRM driver,
export it to a dma-buf file descriptor with DRM_IOCTL_PRIME_HANDLE_TO_FD, then hand
that fd straight to VIDIOC_QBUF as shown above. No pixel data is copied at any point —
only the buffer’s ownership and synchronization move between subsystems.
Enabling Streaming With VIDIOC_STREAMON
Priming the queue only prepares buffers — it does not start the hardware. Streaming must be
explicitly turned on with VIDIOC_STREAMON, which tells the videobuf2 core (the modern
buffer-management layer behind every V4L2 driver) that this queue is now live. Internally, this is
the point where the driver typically starts the underlying hardware pipeline — enabling the sensor,
kicking off the ISP, or arming the DMA engine — and only after this call will
VIDIOC_DQBUF ever return a filled buffer.
static int ep_v4l2_streamon(int fd)
{
enum v4l2_buf_type type = V4L2_BUF_TYPE_VIDEO_CAPTURE;
if (ioctl(fd, VIDIOC_STREAMON, &type) == -1) {
perror("VIDIOC_STREAMON");
return -errno;
}
return 0;
}
Note that VIDIOC_STREAMON and VIDIOC_STREAMOFF both take a pointer to a
plain enum v4l2_buf_type, not a full struct v4l2_buffer — a common mixup
for newcomers to this free linux kernel development course.
Memory Type Comparison For Enqueueing
| Memory Type | Buffer Owner | Fields Set On QBUF | Typical Use Case |
|---|---|---|---|
| V4L2_MEMORY_MMAP | Kernel | type, memory, index | Simple capture apps |
| V4L2_MEMORY_USERPTR | Application | + m.userptr, length | Custom allocators, pre-existing buffers |
| V4L2_MEMORY_DMABUF | Exporter device (GPU/decoder/etc.) | + m.fd | Zero-copy device-to-device pipelines |
Original Demo: ep_v4l2_qstream
Here’s a small, self-contained demo that primes an MMAP-based buffer queue and enables streaming
on any V4L2 capture node. Build it against a current kernel’s linux/videodev2.h header.
// ep_v4l2_qstream.c — prime buffers and start streaming
#include <stdio.h>
#include <string.h>
#include <errno.h>
#include <fcntl.h>
#include <unistd.h>
#include <sys/ioctl.h>
#include <linux/videodev2.h>
#define EP_BUF_COUNT 4
int main(void)
{
int fd = open("/dev/video0", O_RDWR);
if (fd < 0) { perror("open"); return 1; }
/* Assumes VIDIOC_REQBUFS with count = EP_BUF_COUNT, memory = MMAP,
has already been done earlier in your program. */
for (unsigned int i = 0; i < EP_BUF_COUNT; i++) {
struct v4l2_buffer buf;
memset(&buf, 0, sizeof(buf));
buf.type = V4L2_BUF_TYPE_VIDEO_CAPTURE;
buf.memory = V4L2_MEMORY_MMAP;
buf.index = i;
if (ioctl(fd, VIDIOC_QBUF, &buf) == -1) {
perror("VIDIOC_QBUF");
close(fd);
return 1;
}
printf("ep_v4l2_qstream: queued buffer index %u\n", i);
}
enum v4l2_buf_type type = V4L2_BUF_TYPE_VIDEO_CAPTURE;
if (ioctl(fd, VIDIOC_STREAMON, &type) == -1) {
perror("VIDIOC_STREAMON");
close(fd);
return 1;
}
printf("ep_v4l2_qstream: streaming enabled\n");
close(fd);
return 0;
}
Build and run:
$ gcc -o ep_v4l2_qstream ep_v4l2_qstream.c
$ ./ep_v4l2_qstream
Expected output on a working capture node with four buffers requested:
ep_v4l2_qstream: queued buffer index 0
ep_v4l2_qstream: queued buffer index 1
ep_v4l2_qstream: queued buffer index 2
ep_v4l2_qstream: queued buffer index 3
ep_v4l2_qstream: streaming enabled
If any VIDIOC_QBUF call fails with EINVAL, double-check that
buf.memory matches exactly what you passed to VIDIOC_REQBUFS, and that
buf.index is within the buffer count you actually requested.
Common Mistakes And Troubleshooting
- Calling VIDIOC_STREAMON before any buffer is queued: streaming will turn on,
but your first
VIDIOC_DQBUFwill simply block forever, since there’s nothing for hardware to fill. - Mismatched buf.memory: if the value passed to
VIDIOC_QBUFdoesn’t match the memory type used atVIDIOC_REQBUFS, the ioctl fails with-EINVAL. - Touching buffer memory after enqueueing: reading or writing a queued MMAP or USERPTR buffer before it’s dequeued produces undefined results — the DMA engine may be actively writing to it.
- Forgetting the multi-planar array: on multi-planar capture types, omitting
the
v4l2_planearray onbuf.m.planeswill fail the enqueue with-EINVAL. - Passing a v4l2_buffer to VIDIOC_STREAMON/STREAMOFF: these two ioctls expect a
bare
enum v4l2_buf_type, not a full buffer structure.
Best Practices
- Always prime the entire buffer pool before calling
VIDIOC_STREAMONfor smooth, gap-free capture. - Wrap every ioctl call in an
EINTR-retry helper, since signals can interrupt blocking ioctls on a busy system. - Check driver capability flags (from
VIDIOC_QUERYCAP) before assuming DMABUF or multi-planar support is available on a given device. - Never assume enqueue order equals fill order across all drivers — always trust the
indexfield returned byVIDIOC_DQBUF, covered in the next lecture.
Interview Questions
Why does VIDIOC_QBUF lock a buffer’s memory pages?
Because the underlying transfer is DMA-based and DMA engines operate on physical addresses. If the OS swapped a page out mid-transfer, the DMA engine would keep writing to a physical address that no longer holds the intended buffer, silently corrupting memory. Locking (pinning) the pages guarantees the physical address stays valid for the whole transfer.
What’s the practical difference between USERPTR and DMABUF memory types?
USERPTR buffers are plain user-space memory the application allocates itself and hands to the kernel by address on every enqueue — there’s still a mapping step for the DMA engine internally. DMABUF buffers are pre-existing kernel objects shared by file descriptor across subsystems (like DRM/GEM), enabling true zero-copy sharing between two hardware devices without the CPU ever touching the pixel data.
What happens internally when VIDIOC_STREAMON is called?
The videobuf2 core transitions the queue’s state machine into “streaming,” and the driver’s
start_streaming callback runs, which typically programs and enables the underlying
hardware — sensor, ISP, or DMA channel — so that already-queued buffers begin actually being
filled.
Summary And Key Takeaways
Buffer queueing is the handoff point where your application gives up temporary ownership of a
buffer so the driver’s hardware can fill it via DMA. VIDIOC_QBUF locks that buffer’s
memory and its exact fields depend on the memory type — MMAP needs almost nothing extra, USERPTR
needs an address and length, and DMABUF needs a shared file descriptor, often produced by the DRM
PRIME framework. Priming the whole buffer pool before calling VIDIOC_STREAMON keeps
capture smooth, and streaming itself is a simple on/off switch on the queue’s state machine. With
buffers queued and streaming enabled, the next step in this free linux kernel development
course is dequeueing filled buffers — which is exactly where the next lecture picks up.
Frequently Asked Questions
What does VIDIOC_QBUF actually do in V4L2?
It enqueues a previously allocated buffer onto the driver’s active queue, locking its memory pages so hardware DMA can safely fill it.
Do I need to enqueue buffers when using read/write I/O?
No. The read/write I/O method bypasses the queueing mechanism entirely — enqueueing only applies to the streaming I/O methods: MMAP, USERPTR, and DMABUF.
What error does VIDIOC_DQBUF return if I forgot to queue a buffer first?
It returns -EINVAL, since there’s nothing on the queue to dequeue.
Can I call VIDIOC_STREAMON more than once without STREAMOFF?
No — calling it again on an already-streaming queue is a driver-dependent error condition and should be avoided; always pair STREAMON with a matching STREAMOFF before restarting.
Why prime the buffer queue before starting the capture loop?
Priming gives the hardware a deep enough backlog of empty buffers to write into continuously, preventing frame drops caused by the driver having nowhere to DMA new frame data.
Is DMABUF queueing available on every V4L2 driver?
No — check the device’s capability flags from VIDIOC_QUERYCAP and, ideally,
attempt a small test allocation, since DMABUF support depends on both the exporting and importing
drivers.
What is PRIME in the context of DRM and V4L2?
PRIME is the DRM subsystem’s dma-buf sharing layer, built on top of GEM, that lets a GPU or
display buffer be exported as a file descriptor and imported directly into a V4L2 device using
V4L2_MEMORY_DMABUF.
Continue The Free Linux Kernel Development Course
Next up: dequeueing filled buffers with VIDIOC_DQBUF and building the full capture read loop.
Next Lecture Back To Course Index