Files
chore/specs/feat-database-migration.md
ryan f3e9201afd
Chore App Build, Test, and Push Docker Images / build-and-push (push) Failing after 2m19s
feat: update default persistence to MongoDB and enhance configuration options
2026-07-29 16:45:10 -04:00

9.1 KiB

Feature: Migrate Backend Persistence from TinyDB to MongoDB Atlas

Overview

Goal: Replace the TinyDB JSON-file persistence layer in the Flask backend with MongoDB Atlas while keeping the existing API surface, model definitions, and test infrastructure largely unchanged. Introduce a migration script that exports existing TinyDB data into MongoDB.

User Story: As a developer and operator, I want the application to run on a managed, scalable document database so that concurrent access, multi-worker deployments, and production hosting are reliable.

Rules:

  • Backend routes and response shapes must remain unchanged.
  • Python models in backend/models/ keep their from_dict() / to_dict() interface.
  • Every mutation must continue to emit SSE events through the existing broadcaster.
  • The migration must be reversible (keep TinyDB files as a backup) and idempotent.
  • All existing tests must continue to pass; new MongoDB-backed tests must be added.

Data Model Changes

Backend Model

  • No changes to dataclass fields or to_dict() / from_dict() output.
  • The model's id field is stored as MongoDB's _id primary key. The original id field is removed from storage so documents do not contain duplicate values.
  • On read, _id is mapped back to id and the document is wrapped as a TinyDB-compatible Document with doc_id set.
  • created_at and updated_at remain epoch floats.

Frontend Model

  • No changes. The API contract is preserved.

Backend Implementation

1. Database client initialization

  • Create backend/db/mongo_client.py.
  • Do not create a MongoClient at module import; it is created lazily on first access.
  • get_mongo_client() returns a cached process-level singleton so tests, scripts, and Flask requests all share the same client.
  • For Gunicorn multi-worker deployments, backend/gunicorn.conf.py defines a post_fork hook that calls init_mongo_client() so each worker process gets its own client after forking.
  • Fail-fast settings for development:
    MongoClient(
        os.getenv("MONGO_URI"),
        serverSelectionTimeoutMS=5000,
        connectTimeoutMS=5000,
        maxPoolSize=20,
    )
    

2. Collection mapping

  • One MongoDB collection per existing TinyDB file:
    • childrenchildren.json
    • taskstasks.json
    • routinesroutines.json
    • routine_itemsroutine_items.json
    • routine_schedulesroutine_schedules.json
    • routine_extensionsroutine_extensions.json
    • rewardsrewards.json
    • imagesimages.json
    • pending_rewardspending_rewards.json (legacy)
    • pending_confirmationspending_confirmations.json
    • usersusers.json
    • tracking_eventstracking_events.json
    • child_overrideschild_overrides.json
    • chore_scheduleschore_schedules.json
    • task_extensionstask_extensions.json
    • refresh_tokensrefresh_tokens.json
    • push_subscriptionspush_subscriptions.json
    • digest_action_tokensdigest_action_tokens.json

3. LockedTable MongoDB adapter

  • Refactor backend/db/db.py so that each exported *_db object continues to expose the same methods: search(...), get(...), all(), insert(...), insert_multiple(...), update(...), remove(...), truncate().
  • Implement a new MongoLockedTable class that mirrors the existing LockedTable API but delegates to a MongoDB collection.
  • Keep LockedTable (TinyDB) as a fallback selectable by USE_MONGODB=true|false (default true).
  • Store the model id field as MongoDB _id on writes; do not keep a duplicate id field in storage.
  • Restore the model id field from _id on reads and return TinyDB-compatible Document objects with doc_id set so existing callers that rely on doc_id continue to work.
  • Support update(fields, doc_ids=[...]) and callable updater functions (update(lambda rec: ..., cond)) to match TinyDB behavior.
  • Reuse function names where possible to minimize API file changes.

4. Indexes

  • Create indexes on application startup (backend/main.py) and in the pytest session fixture. Index creation is intentionally not done at module import time so the MongoClient remains lazily initialized.
  • Required indexes:
    • Unique primary key is provided by MongoDB _id (mapped from model id); no separate id index is created.
    • user_id — secondary on children, tasks, routines, routine_items, rewards, images, tracking_events, pending_confirmations, chore_schedules, task_extensions, refresh_tokens, push_subscriptions, digest_action_tokens.
    • child_id — secondary on child_overrides, chore_schedules, task_extensions, pending_confirmations, tracking_events.
    • entity_id + entity_type — compound secondary on pending_confirmations, tracking_events, child_overrides.
    • token — unique on refresh_tokens and digest_action_tokens.

5. Migration script

  • Create backend/scripts/migrate_to_mongodb.py.
  • Reads each TinyDB JSON file in data/db/ (or test_data/db/ when DB_ENV=test).
  • Inserts each _default record into the corresponding MongoDB collection, mapping TinyDB integer keys to the record's own id field.
  • Skips records that already exist by id (idempotent).
  • Backs up TinyDB files to data/db/backups/<timestamp>/ before first run.
  • Prints a summary of migrated records per collection.
  • Usage:
    cd backend
    python -m scripts.migrate_to_mongodb [--dry-run]
    

6. Environment variables

  • MONGO_URI — required when USE_MONGODB=true (the default).
  • MONGO_DB_NAME — optional database name; defaults parsed from MONGO_URI or falls back to chore_db/chore_db_test/chore_db_e2e based on DB_ENV/DATA_ENV.
  • USE_MONGODBtrue to use MongoDB backend (default); false to keep TinyDB.

7. Default data initialization

  • backend/db/default.py (initializeImages, createDefaultTasks, createDefaultRewards) must work with the new MongoDB-backed *_db objects. No API changes.

8. Scheduler compatibility

  • Background schedulers (account_deletion_scheduler, digest_scheduler, etc.) read/write through the same *_db objects, so they require no changes once the adapter is in place.

9. Tracking event ordering

  • db/tracking.py adds a hidden monotonic _seq field to tracking events on insert and sorts by (occurred_at, created_at, _seq) descending. This guarantees deterministic ordering when consecutive events share the same timestamp (common in fast tests).

10. Gunicorn / Docker deployment

  • backend/gunicorn.conf.py provides the post_fork hook required by the wiki plan.
  • backend/Dockerfile now loads the config via -c gunicorn.conf.py.

Backend Tests

  • Add mongomock to test dependencies and create backend/tests/test_mongo_adapter.py that verifies CRUD operations against mongomock.
  • Update backend/tests/conftest.py to set USE_MONGODB=true and point MONGO_URI at mongomock for the default test run.
  • Ensure every existing API test still passes with the MongoDB adapter by running pytest tests/.
  • Add integration test script backend/scripts/run_integration_tests.ps1 that starts a local MongoDB Docker container, runs a targeted pytest suite, and tears down the container.
  • Add frontend/.env.test with MongoDB configuration and update Playwright webServer env to pass USE_MONGODB/MONGO_URI to the Flask backend; E2E database is reset via the existing /auth/e2e-seed endpoint.

Frontend Implementation

  • No frontend changes are required. The API contract remains unchanged.

Frontend Tests

  • Run npx playwright test and confirm E2E tests pass against the MongoDB-backed Flask backend with the isolated chore_db_e2e database.

Future Considerations

  • After a production burn-in period, remove the TinyDB fallback and LockedTable code.
  • Add database connection health-check endpoint.
  • Consider MongoDB transactions for multi-document operations (e.g., child deletion cascades).

Acceptance Criteria (Definition of Done)

Backend

  • backend/db/db.py exports MongoDB-backed *_db objects when USE_MONGODB=true.
  • backend/db/mongo_client.py implements lazy client initialization and a Gunicorn-compatible init hook.
  • All existing API endpoints continue to work without route or response changes.
  • backend/scripts/migrate_to_mongodb.py migrates existing TinyDB JSON files idempotently.
  • Required unique and secondary indexes are created on application startup or migration.
  • Unit tests pass with mongomock (pytest tests/).
  • Integration tests pass against a local Docker MongoDB container (requires a running Docker daemon; script verified to start/stop container when daemon is available).

Frontend

  • E2E setup and representative parent-mode smoke tests pass against the isolated chore_db_e2e mongomock database, with the DB reset via /auth/e2e-seed.

Operations

  • MONGO_URI and optional MONGO_DB_NAME environment variables are documented.
  • Rollback instructions exist to switch back to TinyDB by setting USE_MONGODB=false.