A tablet as a second Hyprland monitor, or: 235 ms of good intentions
I have a 13-inch Android tablet that mostly holds PDFs, and a laptop running Hyprland that could use more screen. The obvious thought: make the tablet a real extended display, not a mirror. Wayland, wlroots, hardware encoders on both ends — this should be a solved problem.
It is solved, but not by the software written to solve it.
The purpose-built option
There is a project built for exactly this: a daemon that creates a virtual output on the compositor, captures it, encodes it with VA-API or NVENC, and ships it to a native Android client with reverse touch and stylus input. Rust, GPL-3, active, prebuilt APK, proper multi-distro packaging with signed artifacts. It even has a dedicated Hyprland path that creates the headless output over the compositor's IPC socket and destroys it on shutdown. No root, no kernel module, no screen-share dialog.
It is the right idea, and on my machine the setup worked almost immediately.
Virtual monitor live, hardware H.264 on the iGPU confirmed in the logs, stream
valid under ffprobe. I connected the tablet.
It was unusable. Not marginally — visibly, frustratingly slow.
Measuring instead of guessing
The temptation with "feels slow" is to start turning knobs. I wrote a measurement first, because the two failure modes look identical from the couch and have nothing to do with each other: low frame rate and high latency.
For frame rate, I pulled the daemon's HTTP stream and read frame presentation timestamps:
ffprobe -v error -select_streams v:0 -show_entries frame=pts_time \
-of csv=p=0 sample.ts
The first result was a trap. Ten frames per second, with gaps of almost exactly 103 ms. But the virtual monitor was empty — no windows on it — and compositors do not repaint an output with no damage. I was measuring a static screen and a keepalive timer, not the pipeline.
Mirroring the laptop panel onto the virtual output, with an animation running to guarantee constant damage, gave the honest number: 12.9 fps, at 2544x1800.
Then the useful experiment. I ran the same test at three resolutions:
| Resolution | MB per frame | Delivered |
|---|---|---|
| 2544x1800 | 17.5 | 10.6 fps |
| 1696x1200 | 7.8 | 12.9 fps |
| 1280x904 | 4.4 | 12.9 fps |
Four times less data per frame, no change. Whatever was wrong cost a fixed amount per frame, not per byte.
That killed my leading theory, which I had been fairly confident about. There is
no dmabuf anywhere in that codebase — the capture path is wl_shm, and the
encoder's entry point is push_frame(frame: &[u8]) followed by
copy_from_slice. Every frame crosses CPU memory as a raw byte buffer and gets
copied three or four times before reaching the GPU again. At 2544x1800 that is
17.5 MB per frame, and at 60 fps it would be gigabytes per second of memory
traffic on an integrated GPU already sharing bandwidth with the CPU. A perfectly
good explanation, and not the one that mattered.
Resolution-independence points at a timeout. It was one:
const EVENT_POLL_TIMEOUT_MS: i32 = 100;
That is the timeout on the poll() waiting for the screencopy frame-ready
event. Any frame not already pending when the poll is entered stalls up to
100 ms. Changing it to 4 and rebuilding took the stream from ~11 fps to ~30.
A real bug, worth reporting upstream. It did not fix the problem.
The number that ended it
Frame rate was never the complaint. Latency was, and latency needs a different measurement: something changes on screen at a known instant, and you time when that change emerges from the stream.
I played a fullscreen black video, flipped it to white through the player's IPC
socket at a recorded timestamp, and had ffmpeg watch the stream for the
brightness jump. Six trials.
235 ms, and that is host-side only — capture, encode, mux, HTTP. It excludes the client's own buffering, which the documentation advertises at 40–120 ms.
I trimmed the obvious queues. The encoder's appsrc was holding four frames;
the transport's appsink sixteen. Cutting both bought 6%, inside the noise. The
latency was not in the queues. It was spread through a pipeline that was never
built for low latency: a shm capture with a CPU roundtrip, a CPU colour
conversion, an MPEG-TS mux, and a plain HTTP byte stream.
The project's README advertises 25–40 ms. Measured host-side alone, on hardware it explicitly supports, it was 235.
I want to be fair to it: it does the compositor part well. Rootless virtual outputs over IPC on wlroots, no portal dialog, sensible packaging. If you want a second screen for reference material — documentation, a chat window, something you look at rather than interact with — 235 ms is survivable. As a monitor you drag windows onto and type into, it is not.
The stack that was not designed for this
Game streaming has spent a decade optimising exactly this pipeline, for people who notice single-frame delays. Sunshine on the host, Moonlight on the tablet. It is not built to be a second monitor, but a virtual display plus a streaming server is a second monitor if you squint.
Three things went wrong before it worked, all of them worth writing down.
The AppImage silently used software encoding. It bundles libva 1.23; the system's radeonsi driver is built against VA-API 1.24. libva refuses the mismatch, VAAPI initialisation fails with the gloriously specific "unknown libva error", and the log continues:
Found H.264 encoder: libx264 [software]
No warning that this is a catastrophic downgrade. It would have performed no
better than what I had just abandoned. Deleting the bundled libva so it would
pick up the system one made it hang at the encoder probe instead. The
distribution package, which links system libva, works. The line to check is
h264_vaapi, not libx264.
The newest version was the wrong version. The current stable release has a
reported wlr-screencopy dmabuf regression on AMD with this exact compositor
version, failing with EGL_NOT_INITIALIZED. The previous release is confirmed
working on the same hardware. I installed the older one deliberately. Latest is
a heuristic, not a rule.
Selecting the output by index silently streamed the wrong screen. The
config took output_name = 1, which resolved correctly during the startup probe
and then fell back to monitor 0 when a session actually started. The result was
a mirror of my laptop panel rather than the virtual display — working, plausible,
and completely wrong. Matching by name instead of index fixed it. The log line
that gives it away names the monitor it chose, and it is easy to skim past.
With those sorted: hardware HEVC on the iGPU, capture over dmabuf with no CPU roundtrip, and the same measurement discipline as before — except now the client instruments itself, so I could read the numbers off its own overlay instead of building a rig.
Host processing latency: 4.9 ms. Against 235. Same laptop, same GPU, same compositor, same virtual output. The entire difference is what happens to a frame between the compositor and the encoder.
An hour lost to a firewall
Before any of that worked, the tablet could not reach the host at all. Moonlight said "connection error". Ping worked fine.
The reflex is to debug the application. Instead I started a plain
python3 -m http.server on an unrelated port and tried that from the tablet. Also
blocked. Thirty seconds of work, and it eliminated the entire application as a
suspect.
Then the direction test: the laptop could open a TCP connection to the tablet, but not the reverse, while ICMP passed both ways. That asymmetry rules out client isolation on the access point and points squarely at inbound filtering on the host.
iptables -L showed an INPUT chain with ACCEPT policy and a single jump into a
VPN-managed chain, which was innocent. The rules were somewhere else — native
nftables, which the iptables view does not show:
chain input {
type filter hook input priority filter
policy drop
ct state {established, related} accept
ip protocol icmp accept
tcp dport ssh accept
}
Every symptom, explained. ICMP accepted, so ping worked. Established connections accepted, so anything the laptop initiated worked. Everything else dropped, with a counter quietly climbing into the thousands of packets.
Two rules fixed it, scoped to private address ranges so they travel to other networks without exposing anything publicly:
ip saddr { 10.0.0.0/8, 172.16.0.0/12, 192.168.0.0/16 } \
tcp dport { 47984, 47989, 48010 } accept
ip saddr { 10.0.0.0/8, 172.16.0.0/12, 192.168.0.0/16 } \
udp dport { 47998, 47999, 48000 } accept
I deliberately left the web configuration interface closed. It has no business being reachable from the network when localhost works.
Wired, and why
At 1080p over WiFi the stream was already good: 6 ms network latency, zero dropped frames. Pushing to the tablet's native resolution brought back stutter and a "slow connection" warning, and the obvious reading — not enough bandwidth — was wrong in an interesting way.
The client picks its bitrate automatically from resolution and frame rate. At native resolution and 120 fps it had chosen 88 Mbit/s. The radio was delivering 55. That gap is the warning.
Frame rate was the first mistake: 120 fps on a desktop buys nothing and doubles the requirement. The second was subtler. I reached for resolution as the bandwidth lever, and it is not — at a fixed bitrate cap, bandwidth is that cap regardless of resolution. Resolution only changes how the bits are spread. An idle desktop costs about 1 Mbit/s whatever the cap says, because HEVC drops far below target on static content. The cap bounds motion peaks, nothing else.
The client also turned out to offer no custom resolutions, only fixed presets, of which exactly one matches a 7:5 tablet panel. Every other option is 16:9 and letterboxes. So the virtual display has to match the panel, and bitrate does the bandwidth work.
Then the cable. The instinct is adb reverse, and it cannot work: adb forwards
TCP, and the video, audio and control channels are all UDP. What does work is
USB tethering, which creates a genuine IP link over the cable — the tablet
becomes a gateway, the laptop picks up an address on that subnet, and the client
connects to it like any other host.
It has to be enabled from the tablet's settings. Setting the USB function from the host reports success and changes nothing.
The improvement was not the one I expected:
USB rtt min/avg/max/mdev = 2.078/2.366/2.829/0.266 ms
WiFi rtt min/avg/max/mdev = 6.170/67.571/115.388/50.244 ms
Average latency mattered less than the spread. WiFi swung between 6 and 115 ms. The decoder has to buffer for the worst case it might see, not the average, so jitter that wide sets the floor for how responsive the whole thing can feel. USB holds within 0.75 ms of itself.
Worth noting the cable negotiated USB 2.0. Practical throughput on the tethered interface is somewhere between 150 and 300 Mbit/s, which is far more than needed, but a better cable would raise a ceiling I never reached.
Where it landed
Video stream: 3392x2400 60.02 FPS
Decoder: hevc, low latency
Frames dropped by your network connection: 0.00%
Average network latency: 4 ms (variance: 1 ms)
Host processing latency min/max/average: 10.0/14.5/10.4 ms
Average decoding time: 5.55 ms
Twenty milliseconds of measured pipeline, roughly 27–37 ms glass to glass once frame pacing is counted. Eight megapixels at a locked 60 fps with nothing dropped. Host processing rose from 4.9 ms to 10.4 going from 1080p to native, which is four times the pixels for barely twice the cost.
| Purpose-built | Game-streaming stack | |
|---|---|---|
| Resolution | 2544x1800 | 3392x2400 |
| Frame rate | 30 fps ceiling | 60.02 fps |
| Host latency | 235 ms | 10.4 ms |
| Dropped frames | — | 0.00% |
Twenty-three times better latency, at four times the resolution and twice the frame rate.
What I would tell myself
I picked the first tool on architecture and packaging quality. Clean Rust workspace, signed artifacts, real design documentation, a comparison table where it won every row. All true, and none of it predicted the thing I actually cared about. The stated figure was 25–40 ms. The measured figure, host-side alone, was 235. I should have spent twenty minutes measuring before spending an afternoon configuring.
The measurements that mattered were cheap. Sweeping three resolutions took ten minutes and killed a wrong theory I was confident in. Starting an unrelated HTTP server took thirty seconds and eliminated an entire application as a suspect. The one expensive measurement — the flash-timing rig — was only necessary because that project has no instrumentation of its own. The replacement reports its own latency, which is its own kind of signal about what the authors expected to be asked.
And the winning setup is two tools built for streaming games to a television, pointed at a headless output, over a charging cable.