A standby authenticator, and the backup I decided not to automate
My 2FA codes were the last thing I owned with no second copy. The password manager on the same box gets backed up nightly and verified; the TOTP secrets gating those same accounts lived in exactly one place, a hosted account at a company that could close it, lose it, or fold. Losing that meant losing every second factor at once.
Ente Auth is an end-to-end encrypted authenticator whose server, museum, you can run yourself. So I ran one, as a cold standby: my phone keeps talking to the hosted service, and the self-hosted instance sits there holding a copy for the day that stops working.
The interesting part is not the deployment, which is small. It is the two decisions I got wrong first and the four traps that cost real time.
Three containers, not five
Ente's own compose file runs museum, Postgres, MinIO, a socat shim, and the web bundle. MinIO and socat exist purely for Photos: the authenticator stores its entities in Postgres and never touches object storage. A maintainer confirms you can skip it.
Worth verifying properly rather than trusting, because the failure mode would be
at startup. Reading server/pkg/utils/s3config/s3config.go:
initialize()builds anaws.Configand asession.NewSessionper bucket, and never makes a network call. So no reachable S3 is needed.- But it ends with
viper.Sub("s3").Unmarshal(&config.fileDataConfig), andviper.Subreturns nil for a missing key. Calling a method on that panics. So ans3:key has to exist somewhere. - The image's own baked
configurations/local.yamlhas one, with empty values. That satisfies it.
Which is why my museum.yaml has no s3: block at all, and why that is a
deliberate omission rather than a lucky one. Related: viper.Sub does not consult
AutomaticEnv, so the s3 section is the one part of museum's config that cannot
come from an environment variable.
services:
ente-museum:
image: ghcr.io/ente/server:<commit-sha>
restart: unless-stopped
ports:
- "127.0.0.1:8081:8080"
volumes:
- ./museum.yaml:/museum.yaml:ro
env_file:
- ./ente-auth.env
depends_on:
ente-postgres:
condition: service_healthy
ente-web:
image: ghcr.io/ente/web:<commit-sha>
restart: unless-stopped
ports:
- "127.0.0.1:3003:3003"
environment:
- ENTE_API_ORIGIN=https://auth-api.example.com
ente-postgres:
image: postgres:15.18
restart: unless-stopped
environment:
- POSTGRES_USER=ente
- POSTGRES_DB=ente_db
env_file:
- ./ente-auth.env
volumes:
- ./pgdata:/var/lib/postgresql/data
healthcheck:
test: pg_isready -q -d ente_db -U ente
130 MB of RAM across all three, under a gigabyte of images. The data is TOTP secrets, so it is measured in kilobytes.
Note the image tags. Upstream publishes only latest and per-commit SHAs, no
semver and no per-tag changelog, so pinning means resolving latest to a digest
by hand and reading git log between commits to know what changed. On a
security-relevant service that is a genuine recurring cost, and it is the single
strongest argument against self-hosting this at all.
Secrets as environment overrides
museum reads config through viper with AutomaticEnv, prefix ENTE, and a
replacer turning dots and hyphens into underscores. So db.password becomes
ENTE_DB_PASSWORD, key.encryption becomes ENTE_KEY_ENCRYPTION, and so on.
That means the four secrets live in a mode-600 env file and museum.yaml contains
nothing sensitive, so it can be committed to a config repo verbatim. One file is
shared by museum and Postgres so the database password exists once rather than in
two places that can drift; the cost is that Postgres's environment also carries
museum's encryption keys, which is fine because reading either already requires
root on the host.
key.hash deserves a warning. It is the hashing key for email lookups, not a
signing key, so changing it orphans every existing user row. Treat it like a
database format, not a rotatable credential.
No SMTP, and the login code is in the log
I did not want a mail credential on this box. Ente handles that better than
expected. From server/pkg/utils/email/email.go:
if viper.GetString("smtp.host") == "" {
log.Infof("Skipping sending email to %s: %s", toEmails[0], subject)
and the verification email's subject is built as:
subject := fmt.Sprintf("Verification code: %s", ott)
So with no SMTP configured, the one-time code lands in the container log verbatim and the request returns success. Logging in from a new client is:
docker logs ente-museum --since 10m | grep -i 'Verification code'
That needs shell access, which is the right trade for something touched a few times a year. You would be at a laptop in that scenario anyway, and it is one fewer credential on a publicly reachable host.
The documented alternative, internal.hardcoded-ott, pins a static verification
code for an address. That permanently reduces login to password-only. Not worth it.
The certificate opens a window
Two facts that are individually fine and together are a problem:
- With
internal.adminsempty, museum treats the first user to register as an admin. - A new certificate appears in Certificate Transparency logs within seconds, and scanners follow immediately.
I had measured the second one before, but seeing it against a brand-new hostname was still bracing. Within five minutes of issuance, from the nginx log:
GET /.env Go-http-client
GET /.git/HEAD Go-http-client
GET /?rest_route=/wp/v2/users/ leakix.net scanner
GET /debug/default/view?panel=config leakix.net scanner
GET / axios/1.16.1
So the order matters. Put an IP allow-list in the 443 blocks before enabling the
vhost, register the one account, set disable-registration: true, pin the real
user id into internal.admins, then remove the allow-list. Verify it took by
POSTing a signup request and expecting a 403 rather than trusting the config.
In fairness the window was narrower than it looks, precisely because there is no SMTP: an attacker reaching the endpoint still cannot read the verification code, and museum caps you at 20 wrong attempts per code and 10 active codes an hour. The allow-list is defence in depth, not the thing holding the door shut.
A footnote on allow-lists behind a CDN: allow sees the real client IP only if
you have set_real_ip_from for the CDN's ranges and real_ip_header pointing at
its client-IP header. Get that wrong and you lock out yourself and nobody else. I
also managed to allow-list the wrong address entirely by measuring my public IP
from the wrong machine. The reliable move is to read the denied request out of the
access log rather than guess.
Four traps
ENTE_API_ORIGIN needs a recreate, not a restart. The web image is plain
nginx serving ten prebuilt static apps, one per port. The API origin is not read
at runtime: an entrypoint script seds a placeholder out of the built JavaScript
on first start. Afterwards the placeholder is gone from the container's writable
layer, so docker restart silently keeps the old value. You need
up -d --force-recreate.
museum reflects Origin back itself. Its CORS middleware sets
Access-Control-Allow-Origin to whatever the request sent. So the reverse proxy
must add no CORS headers of its own; two values makes browsers reject the response
outright. Nothing to configure, which is the nice outcome, but very easy to
"fix" into breakage.
The web app cannot import codes. The docs do not say which clients can. The source does: the web auth app has no import route and no reference to one, while the Flutter client has parsers for plain text, Ente's own encrypted export, and a pile of third-party formats. So the browser is for reading codes only, and seeding the standby means using the mobile or desktop client. When you point that client at your instance, the endpoint field wants the API hostname, not the web app's.
Per-tab sessionStorage breaks two pages in a confusing way. Enabling TOTP
and changing your password both call ensureMasterKeyFromSession, which reads the
master key from sessionStorage. That is scoped to a single tab. The auth token,
however, lives in shared localStorage. So if you open /two-factor/setup in a
new tab, it authenticates fine and then throws at the crypto step, surfacing as
a generic "something went wrong".
The server-side signature is unmistakable once you know: POST /users/two-factor/setup
returns 200 with your user id, and then the enable call never arrives at all. The
failure is entirely client-side. Navigate in the tab you logged in with.
The backup I decided not to automate
This is the part I got wrong first.
Ente's CLI can pull authenticator entities down, decrypt them, and write plain
otpauth:// URIs. So the obvious design is a nightly timer: export, encrypt,
push offsite. I built exactly that, complete with the paranoia the job needs.
That paranoia is worth describing, because it is a nice example of a silent-success
API. fetchRemoteAuthenticatorData logs No data to export and returns
nil, nil when it fetches nothing, so the export command writes an empty file and
exits zero. A naive job would ship a correctly-encrypted, plausible-looking,
completely empty archive every night forever. Mine refused to proceed on zero
otpauth:// lines, on any line that was not one, and on a code count lower than
the previous archive's.
Then I deleted the whole thing, before ever running it.
The reason is what account add persists. It stores the account's token, master
key and secret key, encrypted under a device key in a file sitting next to them.
Together those give whoever holds root on that box the ability to decrypt every
second factor you own, with no password. That box is publicly reachable and also
runs my password manager and the age identity for every other backup on it. The
convenience of not typing a command occasionally is not worth handing a
public-facing machine unattended access to all of my 2FA secrets.
So backups are manual, into pass:
# plain-text export from the desktop client, saved under /tmp (tmpfs)
grep -c '^otpauth://' /tmp/codes.txt
pass insert -m ente/auth-codes < /tmp/codes.txt
pass git push
shred -u /tmp/codes.txt
pass show ente/auth-codes | grep -c '^otpauth://' # same count
pass turns out to be a better destination than anything I would have built:
GPG-encrypted to a key that lives on a hardware token, git-versioned so you get
generations for free, and pushed to a git host that is neither my VPS nor the
object storage every other backup on that box depends on. It is the only backup I
have that does not share a failure domain with the others.
I store the plaintext URIs inside pass rather than nesting Ente's own encrypted
export in it. Nesting needs the hardware token and a separate export password,
and a backup you can lose by forgetting a password is worse than one gated on
something you physically hold.
Those URIs are the shared secrets themselves, so the file is equivalent to the
second factor for every account in it. Hence tmpfs and shred, not rm.
The honest cost of this choice: nothing detects staleness. Automation would have caught a forgotten re-export. I traded a real reliability property for a real security one, which is the kind of trade worth naming out loud rather than pretending the chosen option dominates.
Then I turned the website off
A cold standby is reached from a real client. The browser UI is attack surface with no job to do, and the scanners are still knocking. So it is off:
ente-auth-web off # 404 for the web app
ente-auth-web on # serve it again
ente-auth-web status
The obvious implementation is wrong. Do not unlink the vhost: the hostname then falls through to whichever 443 server nginx picks first and serves that host's certificate, so visitors get a TLS name-mismatch warning instead of a clean refusal. Worse, you have to remember to relink it.
Instead the vhost's location / is a one-line include, and the script rewrites
that snippet between a proxy_pass block and return 404, runs nginx -t, and
reloads only if the test passes:
location / {
include /etc/nginx/snippets/ente-auth-web.conf;
}
Keeping the vhost means the certificate and the port-80
/.well-known/acme-challenge/ location both survive, so renewal keeps working
while the app is off. I checked that rather than assuming it: with the app off,
certbot renew --dry-run succeeded for both hostnames, and a token file dropped
into the webroot was still served over plain HTTP.
404 rather than 403, so the hostname discloses nothing about what is behind it. The app container stays running, bound to loopback and unreachable, so turning it back on needs no restart.
Two things live only in the web bundle, so they need it on temporarily:
/two-factor/setup, which is the only route to enabling TOTP, and
/change-password. The desktop client covers the latter from its Account
settings.
A tangent that was not Ente's fault at all
The desktop client refused to delete a code, showing "Authenticate to delete code", with no prompt ever appearing. It reads like the app wants your account password. It does not.
On Linux the client gates destructive actions behind polkit, with an action set to
auth_self, meaning it wants your local user password. The package installs
the policy file correctly. What was missing was a polkit authentication agent,
so polkit had nothing to draw a prompt with, the authorization check returned
not-authorized, and the app fell through to:
if (!result) { showToast(context, infoMessage); return false; }
The string on screen is the failure message, not a request. Nothing typeable would have satisfied it.
The cause was mundane and entirely mine: my compositor config already had a line starting the agent, but the package providing it had never been installed, so the line failed silently at every login. Every polkit escalation on that machine had been broken for weeks and I had not noticed, because nothing I use day to day needs one.
Worth checking ps -eo comm,args | grep -i polkit after any compositor config
change. A missing agent is invisible until something needs it, and then it presents
as an application bug.
Was it worth it?
For the codes, self-hosting barely moves the risk. Ente is end-to-end encrypted, so museum stores TOTP secrets encrypted under keys derived from your password. A server compromise, self-hosted or not, does not hand them over.
What genuinely got worse is that patching is now my job, on a service where "later" is a bad answer, with no semver tags to make it easy. And a host compromise reached through museum lands root next to my password manager's live database. Same service, much worse neighbours.
What got better is that my second factors now exist in three places with different
failure domains, and one of them is a plain-text list of otpauth:// URIs that any
authenticator on earth can import. That last property is the one I actually care
about. It is also, notably, the one the self-hosted server did not provide.
If I were advising someone starting from scratch: do the pass export first. It
closes most of the gap for none of the upkeep. Run the standby only if you want
codes readable on a phone without the vendor, and if you do, put it behind a VPN
or an allow-list rather than the open internet, because a cold standby has no
business being publicly reachable.