Home Blog

A standby authenticator, and the backup I decided not to automate

My 2FA codes were the last thing I owned with no second copy. The password manager on the same box gets backed up nightly and verified; the TOTP secrets gating those same accounts lived in exactly one place, a hosted account at a company that could close it, lose it, or fold. Losing that meant losing every second factor at once.

Ente Auth is an end-to-end encrypted authenticator whose server, museum, you can run yourself. So I ran one, as a cold standby: my phone keeps talking to the hosted service, and the self-hosted instance sits there holding a copy for the day that stops working.

The interesting part is not the deployment, which is small. It is the two decisions I got wrong first and the four traps that cost real time.

Three containers, not five

Ente's own compose file runs museum, Postgres, MinIO, a socat shim, and the web bundle. MinIO and socat exist purely for Photos: the authenticator stores its entities in Postgres and never touches object storage. A maintainer confirms you can skip it.

Worth verifying properly rather than trusting, because the failure mode would be at startup. Reading server/pkg/utils/s3config/s3config.go:

Which is why my museum.yaml has no s3: block at all, and why that is a deliberate omission rather than a lucky one. Related: viper.Sub does not consult AutomaticEnv, so the s3 section is the one part of museum's config that cannot come from an environment variable.

services:
  ente-museum:
    image: ghcr.io/ente/server:<commit-sha>
    restart: unless-stopped
    ports:
      - "127.0.0.1:8081:8080"
    volumes:
      - ./museum.yaml:/museum.yaml:ro
    env_file:
      - ./ente-auth.env
    depends_on:
      ente-postgres:
        condition: service_healthy

  ente-web:
    image: ghcr.io/ente/web:<commit-sha>
    restart: unless-stopped
    ports:
      - "127.0.0.1:3003:3003"
    environment:
      - ENTE_API_ORIGIN=https://auth-api.example.com

  ente-postgres:
    image: postgres:15.18
    restart: unless-stopped
    environment:
      - POSTGRES_USER=ente
      - POSTGRES_DB=ente_db
    env_file:
      - ./ente-auth.env
    volumes:
      - ./pgdata:/var/lib/postgresql/data
    healthcheck:
      test: pg_isready -q -d ente_db -U ente

130 MB of RAM across all three, under a gigabyte of images. The data is TOTP secrets, so it is measured in kilobytes.

Note the image tags. Upstream publishes only latest and per-commit SHAs, no semver and no per-tag changelog, so pinning means resolving latest to a digest by hand and reading git log between commits to know what changed. On a security-relevant service that is a genuine recurring cost, and it is the single strongest argument against self-hosting this at all.

Secrets as environment overrides

museum reads config through viper with AutomaticEnv, prefix ENTE, and a replacer turning dots and hyphens into underscores. So db.password becomes ENTE_DB_PASSWORD, key.encryption becomes ENTE_KEY_ENCRYPTION, and so on.

That means the four secrets live in a mode-600 env file and museum.yaml contains nothing sensitive, so it can be committed to a config repo verbatim. One file is shared by museum and Postgres so the database password exists once rather than in two places that can drift; the cost is that Postgres's environment also carries museum's encryption keys, which is fine because reading either already requires root on the host.

key.hash deserves a warning. It is the hashing key for email lookups, not a signing key, so changing it orphans every existing user row. Treat it like a database format, not a rotatable credential.

No SMTP, and the login code is in the log

I did not want a mail credential on this box. Ente handles that better than expected. From server/pkg/utils/email/email.go:

if viper.GetString("smtp.host") == "" {
    log.Infof("Skipping sending email to %s: %s", toEmails[0], subject)

and the verification email's subject is built as:

subject := fmt.Sprintf("Verification code: %s", ott)

So with no SMTP configured, the one-time code lands in the container log verbatim and the request returns success. Logging in from a new client is:

docker logs ente-museum --since 10m | grep -i 'Verification code'

That needs shell access, which is the right trade for something touched a few times a year. You would be at a laptop in that scenario anyway, and it is one fewer credential on a publicly reachable host.

The documented alternative, internal.hardcoded-ott, pins a static verification code for an address. That permanently reduces login to password-only. Not worth it.

The certificate opens a window

Two facts that are individually fine and together are a problem:

  1. With internal.admins empty, museum treats the first user to register as an admin.
  2. A new certificate appears in Certificate Transparency logs within seconds, and scanners follow immediately.

I had measured the second one before, but seeing it against a brand-new hostname was still bracing. Within five minutes of issuance, from the nginx log:

GET /.env                            Go-http-client
GET /.git/HEAD                       Go-http-client
GET /?rest_route=/wp/v2/users/       leakix.net scanner
GET /debug/default/view?panel=config leakix.net scanner
GET /                                axios/1.16.1

So the order matters. Put an IP allow-list in the 443 blocks before enabling the vhost, register the one account, set disable-registration: true, pin the real user id into internal.admins, then remove the allow-list. Verify it took by POSTing a signup request and expecting a 403 rather than trusting the config.

In fairness the window was narrower than it looks, precisely because there is no SMTP: an attacker reaching the endpoint still cannot read the verification code, and museum caps you at 20 wrong attempts per code and 10 active codes an hour. The allow-list is defence in depth, not the thing holding the door shut.

A footnote on allow-lists behind a CDN: allow sees the real client IP only if you have set_real_ip_from for the CDN's ranges and real_ip_header pointing at its client-IP header. Get that wrong and you lock out yourself and nobody else. I also managed to allow-list the wrong address entirely by measuring my public IP from the wrong machine. The reliable move is to read the denied request out of the access log rather than guess.

Four traps

ENTE_API_ORIGIN needs a recreate, not a restart. The web image is plain nginx serving ten prebuilt static apps, one per port. The API origin is not read at runtime: an entrypoint script seds a placeholder out of the built JavaScript on first start. Afterwards the placeholder is gone from the container's writable layer, so docker restart silently keeps the old value. You need up -d --force-recreate.

museum reflects Origin back itself. Its CORS middleware sets Access-Control-Allow-Origin to whatever the request sent. So the reverse proxy must add no CORS headers of its own; two values makes browsers reject the response outright. Nothing to configure, which is the nice outcome, but very easy to "fix" into breakage.

The web app cannot import codes. The docs do not say which clients can. The source does: the web auth app has no import route and no reference to one, while the Flutter client has parsers for plain text, Ente's own encrypted export, and a pile of third-party formats. So the browser is for reading codes only, and seeding the standby means using the mobile or desktop client. When you point that client at your instance, the endpoint field wants the API hostname, not the web app's.

Per-tab sessionStorage breaks two pages in a confusing way. Enabling TOTP and changing your password both call ensureMasterKeyFromSession, which reads the master key from sessionStorage. That is scoped to a single tab. The auth token, however, lives in shared localStorage. So if you open /two-factor/setup in a new tab, it authenticates fine and then throws at the crypto step, surfacing as a generic "something went wrong".

The server-side signature is unmistakable once you know: POST /users/two-factor/setup returns 200 with your user id, and then the enable call never arrives at all. The failure is entirely client-side. Navigate in the tab you logged in with.

The backup I decided not to automate

This is the part I got wrong first.

Ente's CLI can pull authenticator entities down, decrypt them, and write plain otpauth:// URIs. So the obvious design is a nightly timer: export, encrypt, push offsite. I built exactly that, complete with the paranoia the job needs.

That paranoia is worth describing, because it is a nice example of a silent-success API. fetchRemoteAuthenticatorData logs No data to export and returns nil, nil when it fetches nothing, so the export command writes an empty file and exits zero. A naive job would ship a correctly-encrypted, plausible-looking, completely empty archive every night forever. Mine refused to proceed on zero otpauth:// lines, on any line that was not one, and on a code count lower than the previous archive's.

Then I deleted the whole thing, before ever running it.

The reason is what account add persists. It stores the account's token, master key and secret key, encrypted under a device key in a file sitting next to them. Together those give whoever holds root on that box the ability to decrypt every second factor you own, with no password. That box is publicly reachable and also runs my password manager and the age identity for every other backup on it. The convenience of not typing a command occasionally is not worth handing a public-facing machine unattended access to all of my 2FA secrets.

So backups are manual, into pass:

# plain-text export from the desktop client, saved under /tmp (tmpfs)
grep -c '^otpauth://' /tmp/codes.txt

pass insert -m ente/auth-codes < /tmp/codes.txt
pass git push
shred -u /tmp/codes.txt

pass show ente/auth-codes | grep -c '^otpauth://'   # same count

pass turns out to be a better destination than anything I would have built: GPG-encrypted to a key that lives on a hardware token, git-versioned so you get generations for free, and pushed to a git host that is neither my VPS nor the object storage every other backup on that box depends on. It is the only backup I have that does not share a failure domain with the others.

I store the plaintext URIs inside pass rather than nesting Ente's own encrypted export in it. Nesting needs the hardware token and a separate export password, and a backup you can lose by forgetting a password is worse than one gated on something you physically hold.

Those URIs are the shared secrets themselves, so the file is equivalent to the second factor for every account in it. Hence tmpfs and shred, not rm.

The honest cost of this choice: nothing detects staleness. Automation would have caught a forgotten re-export. I traded a real reliability property for a real security one, which is the kind of trade worth naming out loud rather than pretending the chosen option dominates.

Then I turned the website off

A cold standby is reached from a real client. The browser UI is attack surface with no job to do, and the scanners are still knocking. So it is off:

ente-auth-web off      # 404 for the web app
ente-auth-web on       # serve it again
ente-auth-web status

The obvious implementation is wrong. Do not unlink the vhost: the hostname then falls through to whichever 443 server nginx picks first and serves that host's certificate, so visitors get a TLS name-mismatch warning instead of a clean refusal. Worse, you have to remember to relink it.

Instead the vhost's location / is a one-line include, and the script rewrites that snippet between a proxy_pass block and return 404, runs nginx -t, and reloads only if the test passes:

location / {
    include /etc/nginx/snippets/ente-auth-web.conf;
}

Keeping the vhost means the certificate and the port-80 /.well-known/acme-challenge/ location both survive, so renewal keeps working while the app is off. I checked that rather than assuming it: with the app off, certbot renew --dry-run succeeded for both hostnames, and a token file dropped into the webroot was still served over plain HTTP.

404 rather than 403, so the hostname discloses nothing about what is behind it. The app container stays running, bound to loopback and unreachable, so turning it back on needs no restart.

Two things live only in the web bundle, so they need it on temporarily: /two-factor/setup, which is the only route to enabling TOTP, and /change-password. The desktop client covers the latter from its Account settings.

A tangent that was not Ente's fault at all

The desktop client refused to delete a code, showing "Authenticate to delete code", with no prompt ever appearing. It reads like the app wants your account password. It does not.

On Linux the client gates destructive actions behind polkit, with an action set to auth_self, meaning it wants your local user password. The package installs the policy file correctly. What was missing was a polkit authentication agent, so polkit had nothing to draw a prompt with, the authorization check returned not-authorized, and the app fell through to:

if (!result) { showToast(context, infoMessage); return false; }

The string on screen is the failure message, not a request. Nothing typeable would have satisfied it.

The cause was mundane and entirely mine: my compositor config already had a line starting the agent, but the package providing it had never been installed, so the line failed silently at every login. Every polkit escalation on that machine had been broken for weeks and I had not noticed, because nothing I use day to day needs one.

Worth checking ps -eo comm,args | grep -i polkit after any compositor config change. A missing agent is invisible until something needs it, and then it presents as an application bug.

Was it worth it?

For the codes, self-hosting barely moves the risk. Ente is end-to-end encrypted, so museum stores TOTP secrets encrypted under keys derived from your password. A server compromise, self-hosted or not, does not hand them over.

What genuinely got worse is that patching is now my job, on a service where "later" is a bad answer, with no semver tags to make it easy. And a host compromise reached through museum lands root next to my password manager's live database. Same service, much worse neighbours.

What got better is that my second factors now exist in three places with different failure domains, and one of them is a plain-text list of otpauth:// URIs that any authenticator on earth can import. That last property is the one I actually care about. It is also, notably, the one the self-hosted server did not provide.

If I were advising someone starting from scratch: do the pass export first. It closes most of the gap for none of the upkeep. Run the standby only if you want codes readable on a phone without the vendor, and if you do, put it behind a VPN or an allow-list rather than the open internet, because a cold standby has no business being publicly reachable.