Remote Data Engineer

Reliable data pipelines. Clear business impact.

Python, SQL, PySpark and AWS engineer with corporate-scale experience, business fluency and a portfolio built around the workflows hiring teams actually need: ingestion, orchestration, data quality and reproducible delivery.

Samuel Fersil

GitHub profile

SamHavocH

Brazil, UTC -03:00

20

Public repositories

12

Stars

4

Followers

Data Engineering

Main focus

PythonSQLPySparkAWS GlueAthenaS3QuickSightDockerAirflowdbtPostgreSQLLinux

About

Engineering with business context.

Samuel Fersil builds reliable data systems with Python, Linux, SQL, PySpark, Docker and AWS. His Administration background and Data Science & Analytics studies help him translate business questions into pipelines, data models and operational routines that are clear, testable and useful to decision makers.

Enterprise-ready

Hands-on exposure to large-scale corporate environments, including Itau Unibanco, with the discipline expected in regulated, high-volume contexts.

Built to be operated

Projects emphasize retries, observability, validation, documentation and repeatable local execution, not only happy-path demos.

International fit

Strong fit for remote teams that need clear communication, async ownership and dependable delivery across time zones.

Portfolio

Proof through projects.

The repositories show end-to-end thinking across lakehouse architecture, API ingestion, orchestration, streaming data quality, Linux productivity and automation. They are organized to make technical review fast.

pyspark-medallion-platform

A PySpark lakehouse project that demonstrates how to structure bronze, silver and gold layers, optimize Parquet outputs and run incremental processing with production-style conventions.

PySparkLakehouseParquet
  • Medallion architecture
  • Incremental ETL
  • Distributed processing
View repo

api-to-postgres-pipeline

A practical batch ETL pipeline for API-to-database ingestion, with incremental loads, retries, schema normalization and Dockerized execution for easy technical review.

PythonPostgreSQLDocker
  • API ingestion
  • Retry logic
  • Schema normalization
View repo

airflow-dbt-warehouse

An analytics engineering environment combining Airflow, dbt and PostgreSQL to show orchestration, tested transformations and warehouse modeling in a reproducible local stack.

AirflowdbtWarehouse
  • Bronze/Silver/Gold
  • Testing
  • Docker Compose
View repo

streaming-kafka-data-quality

A streaming data quality project built to demonstrate real-time validation patterns, dead-letter handling and observability concerns that matter in production pipelines.

KafkaStreamingData Quality
  • Observability
  • Dead-letter queues
  • Data validation
View repo

.dotfiles

A Linux and terminal workflow repository that shows environment discipline, reproducible setup habits and productivity choices for engineering work.

LinuxShellAutomation
  • Terminal workflow
  • Setup automation
  • Developer productivity
View repo

validador_documentos

A Python automation project for document quality validation, combining computer vision, business rules, scoring and structured outputs for operational use cases.

PythonValidationAutomation
  • Business rules
  • Quality scoring
  • CSV and JSON reports
View repo

Engineering Depth

Systems fundamentals under pressure.

A technical case that connects Linux internals, debugging, infrastructure fundamentals and English technical communication beyond traditional data engineering projects.

Linux internals

Linux From Scratch: Building Linux from Source

Built a Linux system from source following the Linux From Scratch book during a seven-day live challenge, presented primarily in English. The project covered toolchain bootstrapping, package compilation, chroot environment setup, kernel configuration, bootloader setup, systemd integration and troubleshooting.

  • Built a Linux system from source using the Linux From Scratch book
  • Completed a seven-day live technical challenge
  • Presented the process primarily in English
  • Worked through compilation, configuration and boot issues
  • Strengthened Linux, shell, systems debugging and infrastructure fundamentals
boot log

lfs: $ make tools && enter-chroot

lfs: [kernel] configuring filesystems, drivers and init

lfs: [boot] installing GRUB and validating systemd targets

lfs: [debug] tracing compile, mount and boot failures

7

days

EN

live

LFS

source

LinuxBashGCCMakeKernelGRUBchrootsystemdFilesystemsDebugging

Core Stack

Tools for reliable delivery.

Data Engineering

PythonSQLPySparkETL/ELTData QualityData ModelingBatch ProcessingStreaming

Cloud & AWS

AWS GlueAmazon AthenaAmazon S3AWS QuickSightCloud data pipelines

Analytics Engineering

AirflowdbtBronze/Silver/GoldPostgreSQLData WarehouseTesting

DevOps & Tools

DockerDocker ComposeGitGitHubLinuxShellVS CodeJupyter

GitHub Highlights

Implementation, in public.

The profile currently shows 20 public repositories, 12 stars and a focused portfolio around PySpark, Airflow, dbt, Docker, PostgreSQL, Kafka and Linux. It is structured as a technical signal for recruiters, hiring managers and engineering reviewers.

View GitHub profile

20

Public repositories

12

Stars

4

Followers

Data Engineering

Main focus

Career Narrative

Business to data engineering.

  1. Started with Administration, building a foundation in process, operations and business context

  2. Developed Data Science & Analytics knowledge through MBA-level studies at USP/ESALQ

  3. Gained exposure to corporate-scale environments, including Itau Unibanco

  4. Focused the portfolio on Python, AWS, Linux, PySpark and production-minded data workflows

  5. Now targeting remote international Data Engineer roles where reliability and ownership matter

Need a Data Engineer?

I am open to remote international opportunities involving data pipelines, cloud platforms, automation, ETL/ELT, analytics engineering and data quality.