Home Blog

Dragging a reMarkable 2 forward three years of OS releases

My reMarkable 2 had been running OS 3.2.3.1595 since early 2023, with KOReader launched by a shell script, rm2fb doing the drawing, and a swipe up from the bottom edge toggling between the reader and the reMarkable interface. All of it installed by hand. No Toltec, no package manager, no record of what had been done beyond the files themselves.

It worked, which is the problem with setups like this. There is no reason to touch them until you want one new thing, and then you discover the new thing needs three years of accumulated ecosystem you skipped.

The new thing was reManager, a desktop app for managing mods on reMarkable tablets through the Vellum package manager. This is the story of what it took to get there, including two failures that were interesting and one measurement that surprised me.

What was actually installed

First job was working out what the tablet was running, since I had no notes. ps was more informative than any config file:

651 root  {ko.sh} /bin/sh /home/root/scripts/ko.sh
654 root  rm2fb-server
658 root  {koreader.sh} /bin/sh /home/root/apps/koreader/koreader.sh
663 root  {reader.lua} ./luajit ./reader.lua
241 root  /home/root/apps/touchinjector

The launcher script explains the architecture in fifteen lines:

systemctl stop xochitl
LD_PRELOAD=/home/root/.lib/librm2fb_server.so.1.0.1 /usr/bin/xochitl &
sleep 2
export LD_PRELOAD=/home/root/.lib/librm2fb_client.so.1.0.1
/home/root/apps/koreader/koreader.sh

xochitl is the reMarkable's own application. The trick is to kill it, then restart it under an LD_PRELOAD that turns it into a framebuffer server for someone else. The rM2 has no writable /dev/fb0, so rm2fb fakes one by hooking the one process that does have display access.

The swipe was a separate mechanism. touchinjector is a compiled ARM binary, and strings gives up its whole design:

/dev/input/event2
swipe right
swipe left
execute custom script
~/scripts/swipeup.sh

The script path is compiled in. There is no config file, and no way to point it somewhere else without rebuilding from source I no longer had. Worth knowing before planning any migration around it.

Back up first, and verify the backup

The tablet keeps /home on its own 6.9 GB partition, separate from the two OS slots. That is the part worth preserving, and rsync over the USB network interface pulls it in one pass:

rsync -a --numeric-ids --delete [email protected]:/home/root/ home-root/

811 MB, about a minute. Then the step people skip, which is checking that the copy matches:

rsync -rn --checksum --delete --itemize-changes \
  [email protected]:/home/root/ home-root/

Dry run, checksum comparison, itemized output. Zero differing files. This reported thirteen "skipping non-regular file" lines, all symlinks, which --checksum in dry-run mode does not compare. They were in the real copy. One of them mattered a lot later:

apps/koreader/fonts/CustomFonts -> /home/root/CustomFonts/

That symlink is how KOReader finds custom fonts. Miss it in a restore and 82 MB of fonts silently stop appearing.

The two-stage climb

The plan was to reach 3.27.3.0, the newest image reMarkable publishes for the rM2. Every gesture and launcher package in the Vellum index requires 3.23 or newer, most require 3.26, so almost nothing would install on 3.2.3.

reManager installs a specific version rather than whatever OTA offers, which is what makes this controllable. But the first attempt to reason about it was wrong, and the source explained why:

const minSwupdateVersion = "3.11.2.5"

func (a *App) isLegacyOS() bool {
	v := a.activeOSVersion()
	return v != "" && rmversion.Compare(v, minSwupdateVersion) < 0
}

reMarkable changed updaters partway through the 3.x series. Older systems use update_engine, an Omaha-protocol client; newer ones use swupdate. My tablet had no swupdate binary at all, so reManager correctly classified it as legacy and offered only the older .signed images, which stop at 3.11.2.5. The newer .swu images start at 3.18.1.1. The two ranges do not overlap, so it is two installs with a reboot between them.

The legacy path is nicer than I expected. reManager runs a small Omaha server on the desktop, tunnels it to the tablet over SSH, rewrites update.conf to point at the tunnel, and restores the original before rebooting. Someone had clearly done this to my tablet before, because update.conf still carried the fingerprint:

#SERVER=http://10.11.99.2:8000

The rollback runs out

The rM2 has two OS partitions and updates write to whichever is idle. I had assumed that meant the whole climb was reversible. Checking the partitions directly showed otherwise, because both slots held the same version:

mount -o ro /dev/mmcblk2p3 /tmp/p3chk
cat /tmp/p3chk/etc/version   # 20230227165950

So the slots fill like this:

Point p2 p3 Can return to
Start 3.2.3 (active) 3.2.3 3.2.3
After stage 1 3.2.3 3.11.2.5 (active) 3.2.3
After stage 2 3.27.3.0 (active) 3.11.2.5 3.11.2.5

Stage two overwrites the last copy of the original. After it, the oldest reachable system is 3.11.2.5, because a 3.27 system will not accept a legacy .signed image and the .swu range does not go low enough. Going further back means USB recovery flashing, which erases the device.

That is not a reason to avoid the upgrade. It is a reason to stop between the stages and confirm the tablet is healthy while retreat is still free.

rm2fb was never coming along

rm2fb hooks xochitl at hardcoded addresses, one set per release, listed in src/shared/config.cpp:

!3.2.3.1595
update addr 0x3a1a10
create addr 0x3a4364
...

Forty-five versions are listed. The newest is 3.3.2.1666. Mine was on the list, which is exactly why the tablet worked; anything past 3.3.2 is not, and adding one means disassembling a new xochitl to find five function addresses.

The replacement is not a newer rm2fb. Current KOReader supports three display backends on the rM2, and picks between them by environment variable:

if is_qtfb_native then
    fb_module = "ffi/framebuffer_qtfb"
elseif is_blight_native then
    fb_module = "ffi/framebuffer_blight"
else
    fb_module = "ffi/framebuffer_mxcfb"
end

blight comes from Oxide, qtfb from AppLoad. The check that used to fail hard now returns early:

if is_rm2 then
    if is_qtfb_native or is_blight_native then return Remarkable2 end
    if not os.getenv("RM2FB_SHIM") and not is_qtfb_shimmed then
        error("reMarkable 2 requires a RM2FB server and client to work ...")

This has a consequence for migration: the old KOReader install cannot be carried over. v2025.10 has no blight backend in it at all. Only settings move, and the program comes from the package manager. Which is the correct division anyway, and the reason the migration script I wrote copies exactly six paths: settings/, settings.reader.lua, history.lua, defaults.custom.lua, styletweaks/, and clipboard/.

Reading positions needed nothing. KOReader stores them in .sdr directories next to each book, and /home/root/books never moved. Forty-one of them came through untouched.

Oxide would not start

Both stages went cleanly. Packages installed. KOReader registered itself with exactly the launch entry the packaging promised:

{
    "bin": "/home/root/xovi/exthome/appload/koreader/koreader.sh",
    "environment": { "KO_USE_BLIGHT": "1" },
    "flags": ["nopreload"]
}

And then launcherctl switch-launcher oxide --start failed:

Job for tarnish.service failed because a timeout was exceeded.

The journal was unusually clear about where it died:

tarnish[1407]: Info: Starting wpa_supplicant...
tarnish[1407]: Info: Waiting for wpa_supplicant dbus service...
systemd[1]: tarnish.service: start operation timed out. Terminating.

Ninety seconds of waiting for a D-Bus name that was never going to appear. The reason is a start condition on the supplicant unit:

ExecCondition=bash -c "! grep -q '^\s*wifi\s*=\s*off\s*$' ${CONFIG_DIR}/csl.conf"

My csl.conf contained exactly wifi=off. I keep the radio off on a device whose entire purpose is to not be connected to anything. So systemd skipped wpa_supplicant, nothing claimed fi.w1.wpa_supplicant1, and Oxide's WifiAPI blocked forever on a service that had been deliberately disabled.

Oxide will not start while WiFi is off. Turning the radio on and retrying worked immediately. It is an unfortunate coupling: a launcher should not require a network stack to draw a menu.

Nothing was harmed by the failure, because Oxide ships an OnFailure handler that stops every launcher, starts xochitl, and if even that fails resets the launcher configuration back to none. The tablet stayed usable throughout.

That last sentence is true of a launcher switch you triggered yourself, and I believed it more generally than I should have. It does not hold on boot, as a later reboot demonstrated.

One more gap

With Oxide running, systemctl is-enabled tarnish still reported disabled, and xochitl.service was still in multi-user.target.wants. switch-launcher changes the running state, not the boot state; the enable case in Oxide's launcherctl backend is an empty no-op:

is-enabled | enable | disable)
    ;;

So the switch does not survive a reboot until you do it yourself:

systemctl enable tarnish.service
systemctl disable xochitl.service

The failure handler still works afterwards, because systemctl start operates fine on a disabled unit.

The same bug again, but worse

Everything above was written with the tablet working. Then I rebooted it.

Boot logo, then a blank screen that stayed blank. The tablet answered ping and SSH, so nothing was seriously wrong, but the display showed nothing at all.

tarnish.service          activating   (stuck)
xochitl.service          active
wpa_supplicant.service   inactive (dead, Result: exec-condition)

This is the same wifi=off trap as before. I had turned the radio back on to get Oxide started, tested everything, and then turned it off again, because a device whose purpose is to not be connected to anything should not be connected to anything. csl.conf lives on /home, so the setting survived the reboot, and so did the bug.

Two things about this were worse than the first encounter.

The first is that xochitl reported active while the screen stayed blank. The OnFailure handler had run and had not helped. Whatever guarantee it offers during a manual launcher switch, it did not deliver a usable screen here. I had written that the tablet stays usable throughout, and on boot it does not.

The second is the shape of the fix. Turning WiFi on requires tapping a setting on the screen, and the failure mode is that there is no screen. Without SSH over the USB cable there is no way in. That is a bad place for a daily-driver device to be one setting away from.

So the manual fix is not a fix. The trap needed removing.

Removing the trap

The obvious move is a systemd drop-in that deletes the offending condition. Assigning an empty value resets a directive list, so you can clear all three conditions and put back the two harmless ones:

[Service]
ExecCondition=
ExecCondition=mkdir -p ${CONFIG_DIR}
ExecCondition=touch -a ${CONFIG_DIR}/csl.conf ${CONFIG_DIR}/wifi_networks.conf

I did that, set wifi=off, restarted the unit, and it failed:

Could not read interface wlan0 flags: No such device
nl80211: Driver does not support authentication/association or connect commands
wlan0: Failed to initialize driver interface

wlan0 did not exist. lsmod showed cfg80211 and brcmutil loaded, and brcmfmac absent. Turning WiFi off on this device does not just idle the radio. It unloads the driver, and the interface goes with it.

This is why removing the condition alone makes things worse rather than better. Before, the unit was skipped cleanly. Now it starts, dies on a missing interface, and Restart=on-failure retries it five times in two seconds until systemd gives up:

wpa_supplicant.service: Start request repeated too quickly.

Which leaves the service in failed, the D-Bus name still unclaimed, and the screen still blank. A more confident version of the fix, arrived at faster, would have shipped exactly that.

Loading the driver by hand was more encouraging than expected:

$ modprobe brcmfmac && rfkill list
1: phy1: Wireless LAN
	Soft blocked: yes
	Hard blocked: no

The driver comes up soft blocked. So the interface can exist without the radio transmitting, which is precisely the combination the problem calls for: wpa_supplicant needs a wlan0 to attach to, and I need the radio off. wpa_supplicant starts happily against a blocked interface and logs one line about it:

wpa_supplicant[2668]: rfkill: WLAN soft blocked

The finished drop-in has three moving parts:

[Unit]
StartLimitIntervalSec=0

[Service]
ExecCondition=
ExecCondition=mkdir -p ${CONFIG_DIR}
ExecCondition=touch -a ${CONFIG_DIR}/csl.conf ${CONFIG_DIR}/wifi_networks.conf

ExecStartPre=/bin/sh -c 'modprobe -q brcmfmac; for i in 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15; do if [ -e /sys/class/net/wlan0 ]; then break; fi; sleep 1; done; exit 0'

ExecStartPost=/bin/sh -c 'if grep -q "^[[:space:]]*wifi[[:space:]]*=[[:space:]]*off[[:space:]]*$" ${CONFIG_DIR}/csl.conf; then rfkill block wlan; else rfkill unblock wlan; fi'

RestartSec=3

The condition goes away so the D-Bus name always appears. ExecStartPre makes sure there is an interface to attach to, since the driver may not be loaded. ExecStartPost moves the WiFi switch from the service to the radio, so turning WiFi off in the settings still turns the radio off and still costs nothing, it just no longer takes the launcher down with it.

StartLimitIntervalSec=0 is the small print. The stock unit's Restart=on-failure combined with the default rate limit is what converts a transient failure into a permanent one, and a permanently failed supplicant is the blank screen. Removing the limit means the unit keeps trying instead of giving up in the one state I cannot recover from without a cable.

Rebooting with wifi=off now gives what it should have given all along:

uptime            68 s
csl.conf          wifi=off
wpa_supplicant    active
tarnish           active
systemctl --failed   (empty)

There is a pleasing detail in the end state. Once boot settles, something unloads brcmfmac again, wlan0 disappears, and wpa_supplicant carries on holding the D-Bus name with no interface and no driver underneath it. Oxide is satisfied by a name on a bus. It never actually needed a network.

Which is worth stating plainly, because it is the whole bug: a launcher blocked its own startup on a radio, and the radio was never the point.

A clock that had never once survived a reboot

While in there I checked the date, mostly by accident. The tablet thought it was May 2025, fifteen months behind. The hardware clock was worse:

$ hwclock -r
Fri Jan  2 23:29:23 1970

chronyd is the time daemon, and with the radio off it has never had a source. That part is expected. The 1970 reading is the interesting one, because it means the RTC was never being written, so every boot started from nothing and the clock was reconstructed from whatever the filesystem timestamps implied.

Setting it over the cable takes one line, and the second half is the half that matters:

ssh [email protected] "date -u -s '$(date -u +'%Y-%m-%d %H:%M:%S')'; hwclock -w -u"

date -u -s rather than timedatectl, which refuses to set the clock while an NTP daemon is enabled. hwclock -w -u writes the system time to the RTC, which is the step that makes it persist. It survived the next reboot, which is more than it had managed in a year.

An offline device still keeps time. It just has to be told once, and told in the right place.

The swipe, restored

Oxide's own documentation describes the gesture I wanted:

Swipe up from the bottom of the screen

It opens an application switcher rather than toggling directly, so it is one tap more than the old script. In exchange it is a real task switcher that knows about every registered application, instead of a compiled-in path to one shell script. I consider that a trade in my favor.

Three fonts that default to the wrong weight

With everything working I went to add better fonts, and found something worth writing down on its own.

Variable fonts carry a default instance, used when nothing selects an axis position. For a text face you would expect that default to be Regular. It frequently is not:

f = TTFont("SourceHanSerifSC-VF.ttf")
[(a.axisTag, a.minValue, a.maxValue, a.defaultValue) for a in f['fvar'].axes]
# [('wght', 250.0, 900.0, 250.0)]

250 is ExtraLight. That font had been sitting in my fonts directory for years, and every time I read Chinese in it I was reading hairline strokes on a reflective grey display. Google's Bitter and Noto Sans SC both do the same thing, defaulting to 100.

On a backlit screen this is merely thin. On e-ink, where contrast is already the scarce resource, it is the difference between comfortable and squinting.

The fix is to stop shipping the variable font and instantiate the weights you actually want:

from fontTools.varLib import instancer

f = TTFont(src)
instancer.instantiateVariableFont(f, {"wght": 400}, inplace=True)
f['OS/2'].usWeightClass = 400

Names have to be rewritten too (IDs 1, 2, 4, 6, 16, 17), along with fsSelection and head.macStyle, or the reader will not group the styles into one family and will synthesize a fake bold over your real one.

What went on the tablet, all under the SIL Open Font License:

Font Script Why
LXGW WenKai GB Screen Chinese Kai style, screen-hinted, mainland glyph forms. The obvious pick for long-form Chinese on e-ink.
Source Han Serif SC Chinese Rebuilt at 400 and 700 from the ExtraLight-defaulting original.
Noto Sans SC Chinese Proper multi-weight sans, replacing a single-weight file of uncertain provenance.
Literata English Designed for Google Play Books, so designed for exactly this.
Charis English SIL's readability workhorse, holds up at small sizes.
Bitter English Slab serif drawn for screens; the heavier stems suit e-ink.

Four styles each for the Latin families, two weights for the CJK ones, 124 MB in total. /home had 5.3 GB free, so size was never the constraint. The only thing to be careful about is the root filesystem, which sits at 91 percent full with 21 MB spare. Nothing should ever be installed there.

What I would tell myself in 2023

Nothing here was hard, but three things were only obvious in hindsight.

Verify backups by comparing checksums, not by looking at the file count. The verification took a minute and is the reason the rest of this was relaxed.

Read the tool's source before planning around its behavior. One constant, minSwupdateVersion = "3.11.2.5", turned a one-step plan into a two-step one, and I would have found that out the hard way.

Check the partition you intend to roll back to, rather than assuming an A/B scheme means two different things are present. Mine held two copies of the same release, which quietly made stage two a one-way door.

And if a service hangs for exactly ninety seconds, read the last line it logged before the timeout. It had already told me the answer.

The reboot added two more, and they are the ones I expect to reuse.

A workaround you apply by hand is not a fix, and the difference only shows up when the workaround needs the thing that is broken. Turning WiFi on fixed the launcher perfectly, right up until the failure took away the screen I would have tapped to do it.

And when you disable a condition that was protecting you from a failure, check what the failure actually is first. Removing the wifi=off check turned a cleanly skipped unit into a permanently failed one, which was strictly worse than where I started. The missing driver was two commands away, and I found it only because the first version of the fix broke loudly rather than quietly.