Cross-Platform Python Reporting for Linux and Windows
This article describes how to download Wasabi bucket logs, maintain a lifecycle-aware local mirror, parse the logs with Python, and generate CSV reports, as well as an Excel workbook and a PDF report covering operations, errors, and uploaded/downloaded traffic.
The reporting workflow uses the AWS CLI to synchronize a Wasabi logging bucket or prefix to a local folder. The Python reporter parses the current mirror and rebuilds every report on each run. This approach works with bucket lifecycle expiration. When an old log object expires in Wasabi, the next AWS CLI sync with --delete removes the corresponding local copy.
The reporter preserves the original Wasabi request details and adds reporting fields for:
HTTP method and success state
Decoded object key and request URI
S3 operation, Request ID, requester, User-Agent, Version ID, and TLS version when present
Upload/download classification and transfer bytes
Traffic and operation-level CSV summaries, plus a native Excel workbook (wasabi_report.xlsx) and a PDF report (wasabi_report.pdf) with charts for GET/PUT/DELETE volumes, daily trends, and 403/404/5XX error counts
Requirements
Python 3.9 or later recommended
AWS CLI v2 installed and available in PATH
Wasabi access credentials with permission to list and read the logging bucket/prefix
Network access to the Wasabi service endpoint for the logging region
A lifecycle rule on the logging bucket/prefix if automatic source-log expiration is required
Optional: openpyxl (pip install openpyxl) to build the Excel workbook
Optional: matplotlib (pip install matplotlib) to build the PDF report
If either optional library is missing, the reporter prints a notice, skips that output format, and still writes all CSV files. CSV generation never depends on openpyxl or matplotlib.
Download the Reporter
Download the current Python reporter and save it as wasabi_log_report.py. The script contains the configuration section for the logging bucket, optional prefix, endpoint URL, AWS profile, and region.
Download:
Folder Structure
Linux / macOS
/Users/wos-user/Desktop/Wasabi/
├── logs/
│ └── raw/
├── output/
└── runtime/
This is the root defined for mode L inside get_paths() in the script. Edit that function directly if your Linux/macOS host should use a different root path. There is no separate configuration variable for it.
Windows
C:\Wasabi\
├── logs\
│ └── raw\
├── output\
└── runtime\
The script creates the required directories automatically. The raw directory is a local mirror of the configured Wasabi bucket or prefix. CSV reports are written to output.
Configure Wasabi Access
Configure AWS CLI credentials using a default or named profile.
aws configure
For a named profile:
aws configure --profile wasabi
Use the access key and secret key associated with the Wasabi account/IAM user that can read the logging location. The script uses the Wasabi endpoint configured in its CONFIGURATION section.
Configure the Reporter
Open the script and update the CONFIGURATION section. Configure:
Setting | Purpose |
|---|---|
LOG_BUCKET | Bucket containing the Wasabi bucket-log objects. |
LOG_PREFIX | Optional prefix. Use an empty value ("") when logs are stored at the bucket root. |
ENDPOINT_URL | Wasabi S3 endpoint for the region containing the logging bucket. |
AWS_PROFILE | Optional AWS CLI named profile. Leave this empty to use the default credential resolution. |
AWS_REGION | Optional AWS CLI region argument. |
AWS_ONLY_SHOW_ERRORS | When True, suppresses normal AWS CLI sync output and only prints errors. |
UPLOAD_OPERATIONS / | Sets of operation names classified as UPLOAD or DOWNLOAD traffic. Default to REST.PUT.OBJECT and REST.GET.OBJECT. Extend if your environment logs additional upload/download operation names. |
WRITE_EXCEL_REPORT | Set to False to skip building wasabi_report.xlsx on every run. |
WRITE_PDF_REPORT | Set to False to skip building wasabi_report.pdf on every run. |
MAX_ERROR_RECORDS | Setting for a cap on how many individual failing requests are listed on the Excel workbook’s Error Records sheet. Defaults to 5000. Aggregated error counts elsewhere always cover every record regardless of this cap. |
Bucket Root or Prefix
The same script supports both layouts without an additional runtime parameter.
Logs at bucket root:
s3://logging-bucket/
Logs below a prefix:
s3://logging-bucket/bucket-logs/
Set LOG_PREFIX to an empty value for the root, or to the required prefix such as bucket-logs/. The local raw directory may contain nested folders. The parser scans it recursively.
Synchronization and Lifecycle Cleanup
The reporter invokes AWS CLI sync with --delete and --endpoint-url. The Wasabi source remains authoritative.
aws s3 sync s3://<bucket>/<optional-prefix>/ <local-raw-folder> --delete --endpoint-url https://s3.<region>.wasabisys.com
How --delete Works
--delete does not delete objects from Wasabi when downloading. It deletes local destination files that are no longer present in the source. Use a Wasabi lifecycle rule to expire old log objects. The next synchronization then removes their local copies.
Choose a lifecycle retention period long enough to meet your reporting and troubleshooting requirements. Apply the lifecycle rule only to the dedicated logging bucket or logging prefix that should be expired.
Run the Reporter
Linux / macOS
python3 wasabi_log_report.py L
Windows
python wasabi_log_report.py W
To process an existing local mirror without running AWS CLI sync:
python3 wasabi_log_report.py L --skip-sync
On Windows:
python wasabi_log_report.py W --skip-sync
To skip one or both graphical report formats for a single run (CSV output is unaffected):
python3 wasabi_log_report.py L --no-excel
python3 wasabi_log_report.py L --no-pdf
python3 wasabi_log_report.py L --no-excel --no-pdf
Log Parsing Behavior
The parser reads Wasabi bucket-log records and ignores non-record header/separator lines, such as lines beginning with Record and lines beginning with =. Missing values represented by - remain missing in the CSV rather than being converted into synthetic HTTP values.
Object keys and request URIs are included in both raw and URL-decoded form. Unrecognized records are not silently discarded. They are written to parse_errors.csv with a compact location, reason, and original record.
Upload and Download Traffic
Traffic accounting is operation-aware, so PUT and GET requests are not mixed as they are in a conventional web-server bandwidth report.
Operation | Direction | Bytes Used | Counted When |
|---|---|---|---|
REST.PUT.OBJECT | UPLOAD | ObjectSize | Every matching request |
REST.GET.OBJECT | DOWNLOAD | BytesSent | Every matching request |
Other operations | OTHER | 0 | Not transfer traffic |
Classification is based on the Operation field alone (the UPLOAD_OPERATIONS and DOWNLOAD_OPERATIONS sets in the CONFIGURATION section), not on HTTP status. In practice, a failed PUT typically logs an empty ObjectSize and a failed GET typically logs zero BytesSent, so failed requests naturally contribute little or no traffic without needing a separate status check. Every request, successful or not, still appears in the detailed CSV with its own HttpStatus and Successful values, so failures remain fully visible. Byte-range GET requests (HTTP 206) are counted using BytesSent, which reflects the actual bytes returned for that range rather than the full object size.
CSV Output
CSV Report | Description |
|---|---|
wasabi_log_details.csv | Per-request detail, including Request ID, operation, object, status, errors, client, version, and traffic classification. |
wasabi_traffic_summary.csv | Overall upload/download request and byte totals. |
wasabi_operation_summary.csv | Requests, success/failure counts, traffic, and top errors by S3 operation. |
parse_errors.csv | Records that could not be parsed, with location and reason. |
The daily, per-bucket, top-requester, top-User-Agent, top-object, HTTP-error, and S3-error breakdowns previously produced as separate CSV files are available as charted sheets and pages inside the Excel workbook and PDF report described below, and remain fully derivable from wasabi_log_details.csv for any custom analysis.
Excel and PDF Reports
In addition to the CSV files, the reporter builds a native Excel workbook and a PDF report on every run, controlled by the WRITE_EXCEL_REPORT and WRITE_PDF_REPORT settings (Configure the Reporter) or the --no excel / --no-pdf command-line flags (Run the Reporter).
Excel workbook: wasabi_report.xlsx
This requires openpyxl. The workbook contains six sheets, each with native Excel charts built from the current run's data:
Summary—Headline totals, including total requests, date coverage, requests by verb, 403/404/5XX/other 4XX counts, and upload/download traffic in GiB.
Requests by Verb—A table and bar charts of requests and error counts for GET, PUT, DELETE, HEAD, POST, and OTHER.
Errors—403/404/5XX/other 4XX totals, the full HTTP status code distribution, and the top 25 S3 error codes, each with its own chart.
Errors by Operation—Failed requests broken down by S3 operation and error category, with a stacked bar chart.
Daily Activity—GET/PUT/DELETE and error counts per day (line chart) plus daily upload/download traffic in GiB (stacked bar chart).
Error Records—Individual failed requests (timestamp, verb, operation, status, error code, bucket, key, remote IP, requester, request ID), capped at MAX_ERROR_RECORDS rows. A note points to wasabi_log_details.csv for the complete set when the cap is reached.
PDF report: wasabi_report.pdf
This requires matplotlib. The PDF is a print-ready, one-chart-per-page report intended for sharing outside the reporting toolchain:
Cover page with source bucket/prefix, generation timestamp, date coverage, and headline totals.
Requests by operation verb and 403/404/5XX/other 4XX totals, side by side.
Errors by verb, stacked by error category.
Requests by HTTP status code.
Top 15 S3 error codes by occurrence.
Daily trend: GET/PUT/DELETE and errors on one chart, daily upload/download traffic in GiB on another, sharing a date axis (omitted entirely if no record carries a parseable timestamp).
Detailed CSV Fields
The detailed report includes the original request information plus normalized reporting fields. Source file and source line are not included in the detailed CSV.
BucketOwner, Bucket, TimeRaw, TimestampISO8601, TimestampUTC, RemoteIP, Requester, RequestId, Operation
KeyRaw, KeyDecoded, RequestURI, RequestURIDecoded, HttpMethod, HttpStatus, ErrorCode
BytesSent, ObjectSize, TotalTimeMs, TurnAroundTimeMs, Referrer, UserAgent, TLSVersion, VersionId
TrafficDirection, TrafficBytes, TrafficMiB, TrafficGiB, TrafficTiB, Successful
Automation
Linux can use cron. For example, run the report every hour:
0 * * * * /usr/bin/python3 /Users/wos-user/Desktop/Wasabi/wasabi_log_report.py L >> <path to the log>/report.log 2>&1
On Windows, use Task Scheduler to run python.exe with wasabi_log_report.py W as the arguments. Configure the task to run under the account that owns the AWS CLI credentials/profile.
Scheduling
Avoid overlapping runs. Choose an interval that allows the previous AWS sync and report generation to complete before the next scheduled execution.
Validation and Troubleshooting
Run aws --version and confirm that the AWS CLI is available on the same account that runs the script.
Test AWS credentials by listing the configured logging bucket/prefix with the same endpoint URL.
After the first run, confirm logs exist under logs/raw, and CSV files exist under output.
Review parse_errors.csv. A non-empty file may indicate a new or unexpected Wasabi log format that should be added to the parser.
Compare REST.PUT.OBJECT and REST.GET.OBJECT counts against the detailed CSV when validating traffic totals.
Remember that lifecycle expiration is asynchronous. A lifecycle-eligible object may not disappear at the exact moment it becomes eligible.
If wasabi_report.xlsx or wasabi_report.pdf is missing after a run, check the console output for an install notice for openpyxl or matplotlib. CSV files are still produced regardless.
Operational Recommendations
Use a dedicated bucket or prefix for bucket logs.
Scope lifecycle expiration to the intended logging location.
Use least-privilege credentials for the reporting host.
Protect the local raw logs, CSV output, and the Excel/PDF reports because bucket logs can contain object names, requester identifiers, IP addresses, and User-Agent data. The Error Records sheet and PDF error pages carry this same detail.
Retain parse_errors.csv and investigate new record shapes rather than silently dropping them.
Back up or export CSV, Excel, and PDF reports before lifecycle expiration if historical reporting must extend beyond the raw-log retention window.
Script Download
The Python implementation is distributed separately from this documentation, so the guide remains readable and the script can be updated independently.
Download the current script:
Output examples: