Adding rate limiting to an API is a small piece of work. Deciding what the limits should be is not, and it is a decision that belongs to the people who understand how the product is used rather than only to the people implementing it.
One global limit is the wrong shape
A single limit across all endpoints has to accommodate the most demanding legitimate use, which makes it far too generous for the endpoints that need protection. Ten thousand requests an hour might be reasonable for reading a catalogue and absurd for attempting a password.
Limits should be per endpoint class: authentication, public writes, expensive queries, ordinary reads. Each has a different legitimate ceiling and a different abuse profile.
Partition by the right key
Limiting by IP address punishes everyone behind a corporate NAT or a mobile carrier gateway. Limiting by user account does nothing against an attacker who has not authenticated. Most systems need both: by account where one exists, by IP where it does not, and a separate global ceiling to catch distributed abuse.
The endpoints that need the tightest limits
Login, password reset, OTP verification, coupon validation and anything that sends a message to a person. Each has a specific abuse case: credential stuffing, account enumeration, brute-forcing a six-digit code, discovering valid discount codes, and using your infrastructure to spam someone else at your expense.
These deserve limits an order of magnitude tighter than the rest of the API, and they are usually the last places teams get round to protecting.
Tell the client what happened
Return 429 with a Retry-After header and a clear message. A client that knows when to retry backs off correctly. A client that receives an opaque error frequently retries immediately, which turns a rate limit into a retry storm.
Instrument before you enforce
Run the limiter in observation mode first and look at what real traffic actually does. Teams routinely discover that a legitimate integration partner or their own mobile client sits well above the limit they were about to impose. Finding that in a dashboard is much better than finding it in a support queue.