Extracting data that only appears after interaction
The source site does not present one stable table URL. A user first chooses a country, league and season, then opens the complete-results view. The useful records appear only after that sequence of browser actions.
The project automates that journey and returns a reusable dataset containing dates, competition context, home and away teams, scores, goals, outcomes and the source URL.
The objective is not simply to make Selenium click. It is to leave behind data that is consistent enough to analyse and a failure that is specific enough to investigate.
A narrow workflow with clear boundaries
Separating the page object, scraper and command-line interface keeps website mechanics away from data shaping. A changed selector should not require the export code to be rewritten.
Defensive automation for a dynamic page
Explicit waits
The workflow waits for the condition it actually needs instead of guessing a delay.
DOM stability
Extraction begins after the results view has stopped changing.
Fallback clicks
Alternative interaction paths handle elements that are present but awkward to activate.
Selenium Manager
Drivers are resolved without committing a machine-specific executable path.
Headless execution makes the workflow suitable for CI. If the page fails, the application saves a screenshot and the corresponding HTML so visible state and DOM state can be inspected together.
Deterministic tests without scraping the live site
CI uses a local HTML fixture rather than depending on the public website. This makes the parsing and interaction checks repeatable and avoids turning a third party’s outage or redesign into a false project failure.
pytest, coverage and linting run in GitHub Actions. The live site remains the real execution target, but the build verifies the behaviour the repository owns.
What remains fragile
Browser automation still depends on the source site’s interface and terms of use. A major layout change may require selector updates, and responsible use should include conservative request frequency.
For repeated production collection, a documented API would be preferable. Where none exists, the current design confines that fragility to one page object and keeps the output contract stable.