Dragging a reMarkable 2 forward three years of OS releases
My reMarkable 2 had been running OS 3.2.3.1595 since early 2023, with KOReader
launched by a shell script, rm2fb doing the drawing, and a swipe up from the
bottom edge toggling between the reader and the reMarkable interface. All of it
installed by hand. No Toltec, no package manager, no record of what had been
done beyond the files themselves.
It worked, which is the problem with setups like this. There is no reason to touch them until you want one new thing, and then you discover the new thing needs three years of accumulated ecosystem you skipped.
The new thing was reManager, a desktop app for managing mods on reMarkable tablets through the Vellum package manager. This is the story of what it took to get there, including two failures that were interesting and one measurement that surprised me.
What was actually installed
First job was working out what the tablet was running, since I had no notes.
ps was more informative than any config file:
651 root {ko.sh} /bin/sh /home/root/scripts/ko.sh
654 root rm2fb-server
658 root {koreader.sh} /bin/sh /home/root/apps/koreader/koreader.sh
663 root {reader.lua} ./luajit ./reader.lua
241 root /home/root/apps/touchinjector
The launcher script explains the architecture in fifteen lines:
systemctl stop xochitl
LD_PRELOAD=/home/root/.lib/librm2fb_server.so.1.0.1 /usr/bin/xochitl &
sleep 2
export LD_PRELOAD=/home/root/.lib/librm2fb_client.so.1.0.1
/home/root/apps/koreader/koreader.sh
xochitl is the reMarkable's own application. The trick is to kill it, then
restart it under an LD_PRELOAD that turns it into a framebuffer server for
someone else. The rM2 has no writable /dev/fb0, so rm2fb fakes one by
hooking the one process that does have display access.
The swipe was a separate mechanism. touchinjector is a compiled ARM binary,
and strings gives up its whole design:
/dev/input/event2
swipe right
swipe left
execute custom script
~/scripts/swipeup.sh
The script path is compiled in. There is no config file, and no way to point it somewhere else without rebuilding from source I no longer had. Worth knowing before planning any migration around it.
Back up first, and verify the backup
The tablet keeps /home on its own 6.9 GB partition, separate from the two OS
slots. That is the part worth preserving, and rsync over the USB network
interface pulls it in one pass:
rsync -a --numeric-ids --delete [email protected]:/home/root/ home-root/
811 MB, about a minute. Then the step people skip, which is checking that the copy matches:
rsync -rn --checksum --delete --itemize-changes \
[email protected]:/home/root/ home-root/
Dry run, checksum comparison, itemized output. Zero differing files. This
reported thirteen "skipping non-regular file" lines, all symlinks, which
--checksum in dry-run mode does not compare. They were in the real copy. One
of them mattered a lot later:
apps/koreader/fonts/CustomFonts -> /home/root/CustomFonts/
That symlink is how KOReader finds custom fonts. Miss it in a restore and 82 MB of fonts silently stop appearing.
The two-stage climb
The plan was to reach 3.27.3.0, the newest image reMarkable publishes for the rM2. Every gesture and launcher package in the Vellum index requires 3.23 or newer, most require 3.26, so almost nothing would install on 3.2.3.
reManager installs a specific version rather than whatever OTA offers, which is what makes this controllable. But the first attempt to reason about it was wrong, and the source explained why:
const minSwupdateVersion = "3.11.2.5"
func (a *App) isLegacyOS() bool {
v := a.activeOSVersion()
return v != "" && rmversion.Compare(v, minSwupdateVersion) < 0
}
reMarkable changed updaters partway through the 3.x series. Older systems use
update_engine, an Omaha-protocol client; newer ones use swupdate. My tablet
had no swupdate binary at all, so reManager correctly classified it as legacy
and offered only the older .signed images, which stop at 3.11.2.5. The newer
.swu images start at 3.18.1.1. The two ranges do not overlap, so it is two
installs with a reboot between them.
The legacy path is nicer than I expected. reManager runs a small Omaha server on
the desktop, tunnels it to the tablet over SSH, rewrites update.conf to point
at the tunnel, and restores the original before rebooting. Someone had clearly
done this to my tablet before, because update.conf still carried the
fingerprint:
#SERVER=http://10.11.99.2:8000
The rollback runs out
The rM2 has two OS partitions and updates write to whichever is idle. I had assumed that meant the whole climb was reversible. Checking the partitions directly showed otherwise, because both slots held the same version:
mount -o ro /dev/mmcblk2p3 /tmp/p3chk
cat /tmp/p3chk/etc/version # 20230227165950
So the slots fill like this:
| Point | p2 | p3 | Can return to |
|---|---|---|---|
| Start | 3.2.3 (active) | 3.2.3 | 3.2.3 |
| After stage 1 | 3.2.3 | 3.11.2.5 (active) | 3.2.3 |
| After stage 2 | 3.27.3.0 (active) | 3.11.2.5 | 3.11.2.5 |
Stage two overwrites the last copy of the original. After it, the oldest
reachable system is 3.11.2.5, because a 3.27 system will not accept a legacy
.signed image and the .swu range does not go low enough. Going further back
means USB recovery flashing, which erases the device.
That is not a reason to avoid the upgrade. It is a reason to stop between the stages and confirm the tablet is healthy while retreat is still free.
rm2fb was never coming along
rm2fb hooks xochitl at hardcoded addresses, one set per release, listed in
src/shared/config.cpp:
!3.2.3.1595
update addr 0x3a1a10
create addr 0x3a4364
...
Forty-five versions are listed. The newest is 3.3.2.1666. Mine was on the list,
which is exactly why the tablet worked; anything past 3.3.2 is not, and adding
one means disassembling a new xochitl to find five function addresses.
The replacement is not a newer rm2fb. Current KOReader supports three display
backends on the rM2, and picks between them by environment variable:
if is_qtfb_native then
fb_module = "ffi/framebuffer_qtfb"
elseif is_blight_native then
fb_module = "ffi/framebuffer_blight"
else
fb_module = "ffi/framebuffer_mxcfb"
end
blight comes from Oxide, qtfb from AppLoad. The check that used to fail
hard now returns early:
if is_rm2 then
if is_qtfb_native or is_blight_native then return Remarkable2 end
if not os.getenv("RM2FB_SHIM") and not is_qtfb_shimmed then
error("reMarkable 2 requires a RM2FB server and client to work ...")
This has a consequence for migration: the old KOReader install cannot be
carried over. v2025.10 has no blight backend in it at all. Only settings move,
and the program comes from the package manager. Which is the correct division
anyway, and the reason the migration script I wrote copies exactly six paths:
settings/, settings.reader.lua, history.lua, defaults.custom.lua,
styletweaks/, and clipboard/.
Reading positions needed nothing. KOReader stores them in .sdr directories
next to each book, and /home/root/books never moved. Forty-one of them came
through untouched.
Oxide would not start
Both stages went cleanly. Packages installed. KOReader registered itself with exactly the launch entry the packaging promised:
{
"bin": "/home/root/xovi/exthome/appload/koreader/koreader.sh",
"environment": { "KO_USE_BLIGHT": "1" },
"flags": ["nopreload"]
}
And then launcherctl switch-launcher oxide --start failed:
Job for tarnish.service failed because a timeout was exceeded.
The journal was unusually clear about where it died:
tarnish[1407]: Info: Starting wpa_supplicant...
tarnish[1407]: Info: Waiting for wpa_supplicant dbus service...
systemd[1]: tarnish.service: start operation timed out. Terminating.
Ninety seconds of waiting for a D-Bus name that was never going to appear. The reason is a start condition on the supplicant unit:
ExecCondition=bash -c "! grep -q '^\s*wifi\s*=\s*off\s*$' ${CONFIG_DIR}/csl.conf"
My csl.conf contained exactly wifi=off. I keep the radio off on a device
whose entire purpose is to not be connected to anything. So systemd skipped
wpa_supplicant, nothing claimed fi.w1.wpa_supplicant1, and Oxide's WifiAPI
blocked forever on a service that had been deliberately disabled.
Oxide will not start while WiFi is off. Turning the radio on and retrying worked immediately. It is an unfortunate coupling: a launcher should not require a network stack to draw a menu.
Nothing was harmed by the failure, because Oxide ships an OnFailure handler
that stops every launcher, starts xochitl, and if even that fails resets the
launcher configuration back to none. The tablet stayed usable throughout.
That last sentence is true of a launcher switch you triggered yourself, and I believed it more generally than I should have. It does not hold on boot, as a later reboot demonstrated.
One more gap
With Oxide running, systemctl is-enabled tarnish still reported disabled,
and xochitl.service was still in multi-user.target.wants. switch-launcher
changes the running state, not the boot state; the enable case in Oxide's
launcherctl backend is an empty no-op:
is-enabled | enable | disable)
;;
So the switch does not survive a reboot until you do it yourself:
systemctl enable tarnish.service
systemctl disable xochitl.service
The failure handler still works afterwards, because systemctl start operates
fine on a disabled unit.
The same bug again, but worse
Everything above was written with the tablet working. Then I rebooted it.
Boot logo, then a blank screen that stayed blank. The tablet answered ping and SSH, so nothing was seriously wrong, but the display showed nothing at all.
tarnish.service activating (stuck)
xochitl.service active
wpa_supplicant.service inactive (dead, Result: exec-condition)
This is the same wifi=off trap as before. I had turned the radio back on to
get Oxide started, tested everything, and then turned it off again, because a
device whose purpose is to not be connected to anything should not be connected
to anything. csl.conf lives on /home, so the setting survived the reboot,
and so did the bug.
Two things about this were worse than the first encounter.
The first is that xochitl reported active while the screen stayed blank. The
OnFailure handler had run and had not helped. Whatever guarantee it offers
during a manual launcher switch, it did not deliver a usable screen here. I had
written that the tablet stays usable throughout, and on boot it does not.
The second is the shape of the fix. Turning WiFi on requires tapping a setting on the screen, and the failure mode is that there is no screen. Without SSH over the USB cable there is no way in. That is a bad place for a daily-driver device to be one setting away from.
So the manual fix is not a fix. The trap needed removing.
Removing the trap
The obvious move is a systemd drop-in that deletes the offending condition. Assigning an empty value resets a directive list, so you can clear all three conditions and put back the two harmless ones:
[Service]
ExecCondition=
ExecCondition=mkdir -p ${CONFIG_DIR}
ExecCondition=touch -a ${CONFIG_DIR}/csl.conf ${CONFIG_DIR}/wifi_networks.conf
I did that, set wifi=off, restarted the unit, and it failed:
Could not read interface wlan0 flags: No such device
nl80211: Driver does not support authentication/association or connect commands
wlan0: Failed to initialize driver interface
wlan0 did not exist. lsmod showed cfg80211 and brcmutil loaded, and
brcmfmac absent. Turning WiFi off on this device does not just idle the radio.
It unloads the driver, and the interface goes with it.
This is why removing the condition alone makes things worse rather than better.
Before, the unit was skipped cleanly. Now it starts, dies on a missing
interface, and Restart=on-failure retries it five times in two seconds until
systemd gives up:
wpa_supplicant.service: Start request repeated too quickly.
Which leaves the service in failed, the D-Bus name still unclaimed, and the
screen still blank. A more confident version of the fix, arrived at faster,
would have shipped exactly that.
Loading the driver by hand was more encouraging than expected:
$ modprobe brcmfmac && rfkill list
1: phy1: Wireless LAN
Soft blocked: yes
Hard blocked: no
The driver comes up soft blocked. So the interface can exist without the radio
transmitting, which is precisely the combination the problem calls for:
wpa_supplicant needs a wlan0 to attach to, and I need the radio off.
wpa_supplicant starts happily against a blocked interface and logs one line
about it:
wpa_supplicant[2668]: rfkill: WLAN soft blocked
The finished drop-in has three moving parts:
[Unit]
StartLimitIntervalSec=0
[Service]
ExecCondition=
ExecCondition=mkdir -p ${CONFIG_DIR}
ExecCondition=touch -a ${CONFIG_DIR}/csl.conf ${CONFIG_DIR}/wifi_networks.conf
ExecStartPre=/bin/sh -c 'modprobe -q brcmfmac; for i in 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15; do if [ -e /sys/class/net/wlan0 ]; then break; fi; sleep 1; done; exit 0'
ExecStartPost=/bin/sh -c 'if grep -q "^[[:space:]]*wifi[[:space:]]*=[[:space:]]*off[[:space:]]*$" ${CONFIG_DIR}/csl.conf; then rfkill block wlan; else rfkill unblock wlan; fi'
RestartSec=3
The condition goes away so the D-Bus name always appears. ExecStartPre makes
sure there is an interface to attach to, since the driver may not be loaded.
ExecStartPost moves the WiFi switch from the service to the radio, so turning
WiFi off in the settings still turns the radio off and still costs nothing, it
just no longer takes the launcher down with it.
StartLimitIntervalSec=0 is the small print. The stock unit's
Restart=on-failure combined with the default rate limit is what converts a
transient failure into a permanent one, and a permanently failed supplicant is
the blank screen. Removing the limit means the unit keeps trying instead of
giving up in the one state I cannot recover from without a cable.
Rebooting with wifi=off now gives what it should have given all along:
uptime 68 s
csl.conf wifi=off
wpa_supplicant active
tarnish active
systemctl --failed (empty)
There is a pleasing detail in the end state. Once boot settles, something
unloads brcmfmac again, wlan0 disappears, and wpa_supplicant carries on
holding the D-Bus name with no interface and no driver underneath it. Oxide is
satisfied by a name on a bus. It never actually needed a network.
Which is worth stating plainly, because it is the whole bug: a launcher blocked its own startup on a radio, and the radio was never the point.
A clock that had never once survived a reboot
While in there I checked the date, mostly by accident. The tablet thought it was May 2025, fifteen months behind. The hardware clock was worse:
$ hwclock -r
Fri Jan 2 23:29:23 1970
chronyd is the time daemon, and with the radio off it has never had a source.
That part is expected. The 1970 reading is the interesting one, because it means
the RTC was never being written, so every boot started from nothing and the
clock was reconstructed from whatever the filesystem timestamps implied.
Setting it over the cable takes one line, and the second half is the half that matters:
ssh [email protected] "date -u -s '$(date -u +'%Y-%m-%d %H:%M:%S')'; hwclock -w -u"
date -u -s rather than timedatectl, which refuses to set the clock while an
NTP daemon is enabled. hwclock -w -u writes the system time to the RTC, which
is the step that makes it persist. It survived the next reboot, which is more
than it had managed in a year.
An offline device still keeps time. It just has to be told once, and told in the right place.
The swipe, restored
Oxide's own documentation describes the gesture I wanted:
Swipe up from the bottom of the screen
It opens an application switcher rather than toggling directly, so it is one tap more than the old script. In exchange it is a real task switcher that knows about every registered application, instead of a compiled-in path to one shell script. I consider that a trade in my favor.
Three fonts that default to the wrong weight
With everything working I went to add better fonts, and found something worth writing down on its own.
Variable fonts carry a default instance, used when nothing selects an axis position. For a text face you would expect that default to be Regular. It frequently is not:
f = TTFont("SourceHanSerifSC-VF.ttf")
[(a.axisTag, a.minValue, a.maxValue, a.defaultValue) for a in f['fvar'].axes]
# [('wght', 250.0, 900.0, 250.0)]
250 is ExtraLight. That font had been sitting in my fonts directory for years,
and every time I read Chinese in it I was reading hairline strokes on a
reflective grey display. Google's Bitter and Noto Sans SC both do the same
thing, defaulting to 100.
On a backlit screen this is merely thin. On e-ink, where contrast is already the scarce resource, it is the difference between comfortable and squinting.
The fix is to stop shipping the variable font and instantiate the weights you actually want:
from fontTools.varLib import instancer
f = TTFont(src)
instancer.instantiateVariableFont(f, {"wght": 400}, inplace=True)
f['OS/2'].usWeightClass = 400
Names have to be rewritten too (IDs 1, 2, 4, 6, 16, 17), along with
fsSelection and head.macStyle, or the reader will not group the styles into
one family and will synthesize a fake bold over your real one.
What went on the tablet, all under the SIL Open Font License:
| Font | Script | Why |
|---|---|---|
| LXGW WenKai GB Screen | Chinese | Kai style, screen-hinted, mainland glyph forms. The obvious pick for long-form Chinese on e-ink. |
| Source Han Serif SC | Chinese | Rebuilt at 400 and 700 from the ExtraLight-defaulting original. |
| Noto Sans SC | Chinese | Proper multi-weight sans, replacing a single-weight file of uncertain provenance. |
| Literata | English | Designed for Google Play Books, so designed for exactly this. |
| Charis | English | SIL's readability workhorse, holds up at small sizes. |
| Bitter | English | Slab serif drawn for screens; the heavier stems suit e-ink. |
Four styles each for the Latin families, two weights for the CJK ones, 124 MB
in total. /home had 5.3 GB free, so size was never the constraint. The only
thing to be careful about is the root filesystem, which sits at 91 percent full
with 21 MB spare. Nothing should ever be installed there.
What I would tell myself in 2023
Nothing here was hard, but three things were only obvious in hindsight.
Verify backups by comparing checksums, not by looking at the file count. The verification took a minute and is the reason the rest of this was relaxed.
Read the tool's source before planning around its behavior. One constant,
minSwupdateVersion = "3.11.2.5", turned a one-step plan into a two-step one,
and I would have found that out the hard way.
Check the partition you intend to roll back to, rather than assuming an A/B scheme means two different things are present. Mine held two copies of the same release, which quietly made stage two a one-way door.
And if a service hangs for exactly ninety seconds, read the last line it logged before the timeout. It had already told me the answer.
The reboot added two more, and they are the ones I expect to reuse.
A workaround you apply by hand is not a fix, and the difference only shows up when the workaround needs the thing that is broken. Turning WiFi on fixed the launcher perfectly, right up until the failure took away the screen I would have tapped to do it.
And when you disable a condition that was protecting you from a failure, check
what the failure actually is first. Removing the wifi=off check turned a
cleanly skipped unit into a permanently failed one, which was strictly worse
than where I started. The missing driver was two commands away, and I found it
only because the first version of the fix broke loudly rather than quietly.