Enterprise-ready
Hands-on exposure to large-scale corporate environments, including Itau Unibanco, with the discipline expected in regulated, high-volume contexts.
Python, SQL, PySpark and AWS engineer with corporate-scale experience, business fluency and a portfolio built around the workflows hiring teams actually need: ingestion, orchestration, data quality and reproducible delivery.

GitHub profile
Brazil, UTC -03:00
20
Public repositories
12
Stars
4
Followers
Data Engineering
Main focus
About
Samuel Fersil builds reliable data systems with Python, Linux, SQL, PySpark, Docker and AWS. His Administration background and Data Science & Analytics studies help him translate business questions into pipelines, data models and operational routines that are clear, testable and useful to decision makers.
Hands-on exposure to large-scale corporate environments, including Itau Unibanco, with the discipline expected in regulated, high-volume contexts.
Projects emphasize retries, observability, validation, documentation and repeatable local execution, not only happy-path demos.
Strong fit for remote teams that need clear communication, async ownership and dependable delivery across time zones.
Portfolio
The repositories show end-to-end thinking across lakehouse architecture, API ingestion, orchestration, streaming data quality, Linux productivity and automation. They are organized to make technical review fast.
A PySpark lakehouse project that demonstrates how to structure bronze, silver and gold layers, optimize Parquet outputs and run incremental processing with production-style conventions.
A practical batch ETL pipeline for API-to-database ingestion, with incremental loads, retries, schema normalization and Dockerized execution for easy technical review.
An analytics engineering environment combining Airflow, dbt and PostgreSQL to show orchestration, tested transformations and warehouse modeling in a reproducible local stack.
A streaming data quality project built to demonstrate real-time validation patterns, dead-letter handling and observability concerns that matter in production pipelines.
A Linux and terminal workflow repository that shows environment discipline, reproducible setup habits and productivity choices for engineering work.
A Python automation project for document quality validation, combining computer vision, business rules, scoring and structured outputs for operational use cases.
Engineering Depth
A technical case that connects Linux internals, debugging, infrastructure fundamentals and English technical communication beyond traditional data engineering projects.
Built a Linux system from source following the Linux From Scratch book during a seven-day live challenge, presented primarily in English. The project covered toolchain bootstrapping, package compilation, chroot environment setup, kernel configuration, bootloader setup, systemd integration and troubleshooting.
lfs: $ make tools && enter-chroot
lfs: [kernel] configuring filesystems, drivers and init
lfs: [boot] installing GRUB and validating systemd targets
lfs: [debug] tracing compile, mount and boot failures
7
days
EN
live
LFS
source
Core Stack
GitHub Highlights
The profile currently shows 20 public repositories, 12 stars and a focused portfolio around PySpark, Airflow, dbt, Docker, PostgreSQL, Kafka and Linux. It is structured as a technical signal for recruiters, hiring managers and engineering reviewers.
20
Public repositories
12
Stars
4
Followers
Data Engineering
Main focus
Career Narrative
Started with Administration, building a foundation in process, operations and business context
Developed Data Science & Analytics knowledge through MBA-level studies at USP/ESALQ
Gained exposure to corporate-scale environments, including Itau Unibanco
Focused the portfolio on Python, AWS, Linux, PySpark and production-minded data workflows
Now targeting remote international Data Engineer roles where reliability and ownership matter