Add timeout to network fetches in scrape_me and online scrape_html - #2032
Open
bunlongheng wants to merge 2 commits into
Open
Add timeout to network fetches in scrape_me and online scrape_html#2032bunlongheng wants to merge 2 commits into
bunlongheng wants to merge 2 commits into
Conversation
Both urlopen() and requests.get() calls that fetch a user-supplied URL had no timeout, allowing a slow or hostile server to hang the calling thread indefinitely. Added timeout=30 to both fetch paths.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Both fetch paths that reach out to a remote URL have no timeout, so a server that accepts the connection but never responds hangs the caller forever. scrape_me() does urlopen(Request(url, headers=HEADERS)).read() and the online branch of scrape_html() does requests.get(url=org_url, headers=HEADERS).text, neither of which passes timeout=, and neither urllib nor requests sets a default. I added a 30s timeout to both. It's a small change but it means a slow or hostile host can no longer wedge the calling thread indefinitely. Happy to make the value configurable if you'd rather it live in settings.