0
0
Fork 0
mirror of https://github.com/discourse/discourse.git synced 2026-08-08 17:53:55 +08:00
discourse/migrations
Gerhard Schlager 76a7bd90dc MT: Add a TUI progress reporter and use it in the importer too
This is the next part after the worker pool and reporter changes. I needed
something that can show the progress of more than one step at the same time. The
converter will run steps in parallel later, and the single progress bar we use
today (`ExtendedProgressBar`, on top of `ruby-progressbar`) can only show one
step. For now only one step runs at a time, so you see one live line, and
finished steps and notices scroll up into the terminal history. Once the
scheduler is done, showing several lines at once works without more changes.

I first checked which library to use in a small POC: ruby-progressbar (what we
use today), bubbletea-ruby, tty-progressbar, and a hand-written ANSI renderer.
The gems all needed me to fight or patch their internals to get what we need.
bubbletea came closest, but its event loop blocks the other threads while it
waits for input, so the progress updates dropped from about 200/s to 9/s — that
would slow down the actual conversion, not just the display. Working around that
means patching the gem, and keeping notices and finished steps in the terminal
history needs changing how it draws. So I'd be writing the hard parts myself
anyway, on top of behavior I don't control and a heavy dependency. The
hand-written renderer does everything we need, uses much less CPU, and only adds
one small pure-Ruby gem (`unicode-display_width`). It also lets us drop
`ruby-progressbar`.

What it does:

- Shows a live line per running step, with finished steps and notices scrolling
  into the history. Each step ends as done, interrupted (Ctrl-C) or failed.
- Falls back to plain line output for pipes, CI and dumb terminals.
- The importer reports through the same thing now, so the converter and the
  importer share one progress display.

Roughly how a run looks:

    ✓ Converting categories                4,281   0:01
    ✓ Converting users                   312,440   0:03   ⚠ 17 warnings
    ⠋ Converting posts          41%    1,248,776   4:07   ETA 5:52   73,143/s
    ⠋ Converting tags           11%       10,944   0:03   ETA 0:24

The output in pipes and CI looks different than before.
2026-06-25 17:46:04 +02:00
..
bin MT: Split the migrations tooling into separate gems (#40492) 2026-06-02 22:20:03 +02:00
converters MT: Clarify ForkManager hook API names 2026-06-24 13:06:22 +02:00
core MT: Add a TUI progress reporter and use it in the importer too 2026-06-25 17:46:04 +02:00
docs MT: Add disco check — a single entrypoint for all schema and converter checks 2026-06-11 21:27:17 +02:00
importer MT: Add a TUI progress reporter and use it in the importer too 2026-06-25 17:46:04 +02:00
tooling MT: Add a TUI progress reporter and use it in the importer too 2026-06-25 17:46:04 +02:00
.gitignore MT: Split the migrations tooling into separate gems (#40492) 2026-06-02 22:20:03 +02:00
.reek.yml MT: Refactor schema configuration from YAML to Ruby DSL 2026-03-19 18:10:26 +01:00
.rubocop.yml MT: Add a TUI progress reporter and use it in the importer too 2026-06-25 17:46:04 +02:00
AGENTS.md MT: Add disco check — a single entrypoint for all schema and converter checks 2026-06-11 21:27:17 +02:00
CLAUDE.md MT: Refactor schema configuration from YAML to Ruby DSL 2026-03-19 18:10:26 +01:00
README.md MT: Discard inherited source DB connections in worker processes 2026-06-14 20:53:33 +02:00

Migrations Tooling

The migrations/ directory is split into four path-referenced gems:

  • core/Migrations::*: CLI framework, UI, SQLite schemas, DB infrastructure, IntermediateDB models, and the conversion framework (Migrations::Conversion::*).
  • tooling/Migrations::Tooling::*: the schema DSL, disco schema commands, benchmarks.
  • converters/Migrations::Converters::*: public converter implementations + source adapters.
  • importer/Migrations::Importer::*: the row importer and the uploads importer.

All four are wired into the root Gemfile via path: in the optional :migrations group.

Command line interface

The single binary is migrations/bin/disco (commands register dynamically via Migrations::CLI::Registry). Run it without arguments — or with --help — for the authoritative, always-current list of commands:

migrations/bin/disco --help

Rails is booted lazily: only commands that declare requires_rails! (import, upload, schema) load the Discourse app.

Converters

Public converters live in converters/lib/migrations/converters/. To run a private (closed-source) converter, put its code in a subdirectory of private/converters/ (or point MIGRATIONS_PRIVATE_CONVERTERS_PATH at it).

Source DB adapters and fork safety

Worker processes inherit the source DB connection's socket from the main process. Whether that's dangerous depends on the client library: a destructor that only closes the file descriptor is harmless (the parent still holds it, so the kernel sends nothing over the wire), but a destructor that writes a protocol goodbye kills the parent's session as soon as a worker exits — libpq sends a Terminate message, MySQL clients send COM_QUIT.

Adapter::Postgres handles this by registering a ForkManager.after_fork_child hook that calls discard! in each worker: the inherited socket is redirected to /dev/null, and any later use of the adapter in the worker raises DiscardedError. New adapters should follow the same pattern. The discard mechanism itself is library-specific — mysql2 has automatic_close = false, trilogy has a native discard!. To check whether a library needs one at all: connect, fork an empty child that exits normally, wait for it, and query again from the parent (see the fork-safety specs in postgres_spec.rb).

Schema DSL

The schema DSL lives in migrations/tooling/lib/migrations/tooling/schema/dsl/. Config sources are in migrations/tooling/config/schema/. Generated artifacts (SQL, models, enums) are written into migrations/core/.

Key files:

  • table_builder.rb - DSL for defining table configs
  • schema_resolver.rb - Resolves DSL config + DB introspection into final schema
  • conventions_builder.rb - Global column conventions (renames, type overrides)
  • generator.rb - Generates SQL, models, and enums from resolved schema
  • validator.rb - Validates DSL config
  • resolved_schema_validator.rb - Validates resolved schema before generation

Development

Installing gems

bundle config set --local with migrations
bundle install

Updating gems

bundle update --group migrations

Running tests

Each gem has an isolated, no-Rails suite, run from the gem directory:

cd migrations/core       && bundle exec rspec
cd migrations/tooling    && bundle exec rspec
cd migrations/converters && bundle exec rspec
cd migrations/importer   && bundle exec rspec

Specs that need a booted Rails environment are tagged :rails. They are excluded by default and run from the host app's bundle:

cd migrations/<gem> && BUNDLE_GEMFILE=../../Gemfile MIGRATIONS_RAILS=1 bundle exec rspec --tag rails

Linting

bin/lint path/to/file
bin/lint --fix path/to/file

Uses both rubocop and syntax_tree. Always lint changed files.

Known issues

  • Parallel step items must be hashes (or other non-scalar JSON values). Worker processes receive their items as an Oj stream over a pipe, and Oj's stream parser can only detect the end of a bare scalar (a number or string) once the next byte arrives — so a step whose items yields scalars stalls the worker pipeline. Wrap scalar items in a hash (e.g. { id: ... }). This should go away if we switch the worker pipes from Oj to JSON.