Back to project archive

06 · Browser automation · Structured extraction

Football Results Browser Automation

A Selenium application that turns a dynamic, filter-driven football website into clean CSV or JSON without hiding the reliability work behind the browser.

2portable output formats
0live-site calls made by CI tests
8+fields retained for every match
HTML + PNGevidence captured after failures

Extracting data that only appears after interaction

The source site does not present one stable table URL. A user first chooses a country, league and season, then opens the complete-results view. The useful records appear only after that sequence of browser actions.

The project automates that journey and returns a reusable dataset containing dates, competition context, home and away teams, scores, goals, outcomes and the source URL.

The objective is not simply to make Selenium click. It is to leave behind data that is consistent enough to analyse and a failure that is specific enough to investigate.

A narrow workflow with clear boundaries

01
Configure the requestCountry, league, season, output path and format are explicit inputs.
INPUT
02
Navigate through the page objectSelectors and browser actions remain outside the extraction logic.
BROWSE
03
Wait for the required stateExplicit conditions replace machine-dependent fixed sleeps.
SYNC
04
Parse and deduplicateRendered match rows become a consistent record structure.
SHAPE
05
ExportThe same records can be written as CSV or JSON.
DELIVER

Separating the page object, scraper and command-line interface keeps website mechanics away from data shaping. A changed selector should not require the export code to be rewritten.

Defensive automation for a dynamic page

CHOICE 01

Explicit waits

The workflow waits for the condition it actually needs instead of guessing a delay.

CHOICE 02

DOM stability

Extraction begins after the results view has stopped changing.

CHOICE 03

Fallback clicks

Alternative interaction paths handle elements that are present but awkward to activate.

CHOICE 04

Selenium Manager

Drivers are resolved without committing a machine-specific executable path.

Headless execution makes the workflow suitable for CI. If the page fails, the application saves a screenshot and the corresponding HTML so visible state and DOM state can be inspected together.

Deterministic tests without scraping the live site

CI uses a local HTML fixture rather than depending on the public website. This makes the parsing and interaction checks repeatable and avoids turning a third party’s outage or redesign into a false project failure.

pytest, coverage and linting run in GitHub Actions. The live site remains the real execution target, but the build verifies the behaviour the repository owns.

PythonSelenium 4pytestPage objectExplicit waitsCSVJSONGitHub Actions

What remains fragile

Browser automation still depends on the source site’s interface and terms of use. A major layout change may require selector updates, and responsible use should include conservative request frequency.

For repeated production collection, a documented API would be preferable. Where none exists, the current design confines that fragility to one page object and keeps the output contract stable.

Next case studyJenkins Selenium Delivery Pipeline