mirror of
https://github.com/discourse/discourse.git
synced 2026-08-06 07:23:30 +08:00
GitHub oneboxes and the discourse-github plugin talked to GitHub's REST
and
GraphQL API with no rate-limit awareness. On busy instances this
exhausted
GitHub's limits (60 requests/hour unauthenticated, 5000 authenticated),
and
because there was no backoff every render kept hitting GitHub and
re-failing
-- which GitHub's docs warn can get an integration banned. The recently
added PR-status onebox multiplied the number of calls and made it far
worse.
GitHub access was also fragmented: the core onebox engines used OpenURI,
the
discourse-github plugin used Octokit, and the discourse-ai bot tools
used
FinalDestination::HTTP -- three HTTP stacks, three tokens, and
inconsistent
(or entirely missing) error and rate-limit handling.
This introduces a single client, Discourse::GithubApi, that every GitHub
data-API request now flows through. It is built on Faraday with the
SSRF-safe
FinalDestination adapter and:
- authenticates per token (Bearer) and returns plain string-keyed Hashes
(get/post) or raw bodies (raw_get) -- one response shape, no
Octokit/Sawyer
- only ever sends the access token to api.github.com and
raw.githubusercontent.com, rejecting any other absolute URL, so a
user-derived path can never leak a token to an arbitrary host
- backs off on rate limits both reactively (403/429) and proactively
(when
X-RateLimit-Remaining hits 0), honouring Retry-After /
X-RateLimit-Reset,
via a shared Redis flag (GithubRateLimit) keyed per token so each
token's
budget and the shared unauthenticated/IP budget back off independently
- short-circuits while backing off without ever sleeping, so onebox
rendering
and post baking degrade to a plain link instead of blocking a request
- caches ETags and sends If-None-Match, so unchanged resources return
304s
that do not count against the rate limit
Every caller was moved onto it:
- the 6 core GitHub onebox engines, via a slimmed
Onebox::Mixins::GithubApi
adapter that keeps their public methods and translates client errors
back
to the OpenURI::HTTPError vocabulary they already rescue (engines
unchanged)
- the github_blob raw.githubusercontent.com fetch
- the discourse-github plugin (badges, linkback, permalinks, token
validator),
which no longer uses the octokit and sawyer gems (they stay in the
Gemfile for
the discourse-code-review official plugin, which still depends on them)
- the discourse-ai bot's GitHub tools (search code, diff, file content,
search files)
Also adds a GithubOneboxBackoff admin problem check that surfaces while
one of
the onebox token identities is backing off -- scoped to the tokens
resolved by
Onebox::GithubAccess (each configured github_onebox_access_tokens entry
plus the
unauthenticated client) so a backoff on the AI bot or linkback token is
not
misattributed to onebox. Its message points admins at the relevant
setting with
the {{setting:...}} link marker, which problem-check messages now expand
too.
Onebox token resolution is centralised in Onebox::GithubAccess, and the
onebox
cache TTL for transient GitHub failures is shortened so they recover
quickly.
GitHub OAuth login, theme git-clone, the inbound webhook, and the
Oneboxer
FinalDestination URL-resolution special-cases for github.com are
intentionally
out of scope -- they are different concerns, not the rate-limited data
API.
66 lines
2.2 KiB
Ruby
Vendored
66 lines
2.2 KiB
Ruby
Vendored
# frozen_string_literal: true
|
|
|
|
RSpec.describe GithubRateLimit do
|
|
let(:token) { "gh_token_123" }
|
|
|
|
def ttl(token = nil)
|
|
Discourse.redis.without_namespace.ttl(described_class.key(token))
|
|
end
|
|
|
|
describe ".key" do
|
|
it "is per-token, and a shared bucket for unauthenticated requests" do
|
|
expect(described_class.key(token)).to eq(
|
|
"onebox_github_backoff_#{Digest::SHA1.hexdigest(token)}",
|
|
)
|
|
expect(described_class.key(nil)).to eq("onebox_github_backoff_unauthenticated")
|
|
expect(described_class.key("")).to eq("onebox_github_backoff_unauthenticated")
|
|
end
|
|
end
|
|
|
|
describe ".note_rate_limit" do
|
|
it "backs off for the Retry-After duration" do
|
|
described_class.note_rate_limit(token:, retry_after: 90)
|
|
expect(described_class.backing_off?(token)).to eq(true)
|
|
expect(ttl(token)).to be_between(1, 90)
|
|
end
|
|
|
|
it "backs off until x-ratelimit-reset when the remaining budget is 0" do
|
|
described_class.note_rate_limit(token:, remaining: "0", reset_at: 10.minutes.from_now.to_i)
|
|
expect(ttl(token)).to be_between(1, 600)
|
|
end
|
|
|
|
it "does nothing when there is budget left and no Retry-After" do
|
|
described_class.note_rate_limit(token:, remaining: "57")
|
|
expect(described_class.backing_off?(token)).to eq(false)
|
|
end
|
|
|
|
it "clamps the backoff to one hour" do
|
|
described_class.note_rate_limit(token:, remaining: "0", reset_at: 10.days.from_now.to_i)
|
|
expect(ttl(token)).to be <= 1.hour.to_i
|
|
end
|
|
|
|
it "is scoped per token" do
|
|
described_class.note_rate_limit(token:, retry_after: 90)
|
|
expect(described_class.backing_off?("other_token")).to eq(false)
|
|
expect(described_class.backing_off?(nil)).to eq(false)
|
|
end
|
|
end
|
|
|
|
describe ".note_response_headers" do
|
|
it "reads the rate-limit headers regardless of casing" do
|
|
described_class.note_response_headers(
|
|
{ "Retry-After" => "120", "X-RateLimit-Remaining" => "0" },
|
|
token:,
|
|
)
|
|
expect(ttl(token)).to be_between(1, 120)
|
|
end
|
|
end
|
|
|
|
describe ".active?" do
|
|
it "is true while any backoff key exists" do
|
|
expect(described_class.active?).to eq(false)
|
|
described_class.note_rate_limit(token:, retry_after: 30)
|
|
expect(described_class.active?).to eq(true)
|
|
end
|
|
end
|
|
end
|