Data Engineer Resume
0/5 (Submit Your Rating)
CaliforniA
SUMMARY
- Around 6+ years of experience in Big Data Technologies and Software development.
- Extensive experience building data pipelines for processing and migration of data from on premise ETLs to Google Cloud Platform (GCP) using cloud native tools such as BIG query, Cloud Data Proc, Google Cloud Storage, Composer.
- Practical understanding of the Data modelling (Dimensional & Relational) concepts like Star - Schema Modelling, Snowflake Schema Modelling, Fact and Dimension tables.
- Strong knowledge of the Big Data ecosystem and its associated components.
- Architected a cloud native data warehouse on Snowflake. Developed the schema, underlying data architecture, warehouse sizing and capacity requirements. Worked with several stakeholders to identify requirements and deliver.
- Experience in developing applications as well as administering and troubleshooting real-time data processing and ingestion tools - Kafka, NiFi and Spark.
- Strong experience using Spark RDD API, Spark Dataframe/Dataset API, and Spark-SQL frameworks for building end to end data pipelines.
- Strong knowledge of Spark Internals and job execution lifecycle of spark applications.
- Experience building Hadoop clusters on AWS.
- Experience developing Spark jobs using Python.
- Strong experience working with Hive for performing various data analysis.
- Experience with performance tuning of Hive and Spark jobs.
- Good experience in automating end to end data pipelines using Oozie workflow orchestrator and airflow.
- Experience in developing rich interactive Tableau visualizations using various visualizations like Heat and Tree Maps, Bubble Charts, Reference Lines, Dual Axes, Line diagrams, Sankey charts and Geographic Visualizations
- Hands on Bash scripting experience and building data pipelines on Unix/Linux systems.
- Diverse experience in all phases of software development life cycle (SDLC) especially in Analysis, Design, Development, Testing and Deploying of applications.
- Knowledge in various file formats in HDFS like Avro, orc, parquet.
- Excellent communication and interpersonal skills and capable of learning new technologies very quickly
- Experience building serverless web applications using AWS Cognito, API gateway, Lambda, DynamoDB, S3, CloudFormation, and RedShift.
- Experience working on Jenkins, Git, Ansible and other CI/CD tools in Devops.
- Proficient in communicating with people at all levels of hierarchy in the organization and a team player with strong analytical, programming, administration, and problem-solving skills.
TECHNICAL SKILLS
Big Data Technologies: Hadoop (HDFS & MapReduce), Spark, Hive, Sqoop, Oozie, AWS, Kafka, Snowflake, Zookeeper, Airflow, Jenkins, Git, ELK, Redis, Docker and HBase.
Languages /Scripting: Python and Shell Scripting
Databases and Tools: Oracle, MySQL, SQL, NoSQL, PostgreSQL
Platforms: Windows, Linux
IDEs: Pycharm, Spyder and Anaconda
Scheduling: Airflow, Oozie
PROFESSIONAL EXPERIENCE
Confidential, California
Data Engineer
Responsibilities:
- Installed/Configured/Maintained Apache Hadoop clusters for application development based on the requirements using Amazon EKS.
- Involved in converting Hive/SQL queries into Spark transformations using Spark RDDs, Python.
- Optimized Spark applications to improve performance and utilize cluster resources efficiently.
- Developed Spark-Streaming applications to consume the data and for data ingestion.
- Effectively used Sqoop to import data from MySQL, Exadata into HDFS raw layer, perform transformations and publish the data into parquet format.
- Developed a python module to process high volume data in Kafka messaging queues. Built Asynchronous python modules to analyse and modify the data in Kafka streams.
- Developed a Linux style cross-platform command line tool using Python to load 100+ TB data into Snowflake using multi-threading.
- Configurable with defined parameters specifying schema, table, warehouse through the shell or through a config file. Inbuilt resiliency, logging, auditing, and monitoring mechanism for the data being loaded.
- Integrated CloudWatch logs to store all user actions for future purposes and used AWS Athena to analyse data stored in S3 buckets.
- Implemented authentication and authorization using AWS Cognito User Pools, AWS Cognito Identity Pools, Microsoft AD and Lambdas.
- Utilized AWS services like EMR, S3, Glue and Athena extensively for building the data applications.
- Developed Lambda functions in Python for reading data from Kinesis streams and built Dynamo-DB tables using Policies and Roles on the AWS console.
- Participated in architecture discussions and provided solutions for building use cases on AWS.
- Lead a team of data engineers to build and operationalize data pipelines.
Confidential
Data Engineer
Responsibilities:
- Migrating an entire oracle database to BigQuery and using of Tableau for reporting.
- Build data pipelines in airflow in GCP for ETL related jobs using different airflow operators.
- Experience in GCP Dataproc, GCS, Cloud functions, BigQuery.
- Experience in moving data between GCP and Azure using Azure Data Factory.
- Experience in building power bi reports on Azure Analysis services for better performance.
- Used cloud shell SDK in GCP to configure the services Data Proc, Storage, BigQuery Coordinated with team and Developed framework to generate Daily adhoc reports and Extracts from enterprise data from BigQuery.
- Designed and Co-ordinated with Data Science team in implementing Advanced Analytical Models in Hadoop Cluster over large Datasets.
- Wrote scripts in Hive SQL for creating complex tables with high performance metrics like partitioning, clustering, and skewing
- Work related to downloading BigQuery data into pandas or Spark data frames for advanced ETL capabilities.
- Worked with google data catalog and other google cloud APIs for monitoring, query and billing related analysis for BigQuery usage.
- Worked on creating POC for utilizing the ML models and Cloud ML for table Quality Analysis for the batch process.
- Knowledge about cloud dataflow and Apache beam.
- Good knowledge in using cloud shell for various tasks and deploying services.
- Created BigQuery authorized views for row level security or exposing the data to other teams.
- Expertise in designing and deployment of Hadoop cluster and different Big Data analytic tools including Hive, SQOOP, Apache Spark, with Cloudera Distribution.
Confidential
Data Engineer
Responsibilities:
- Developed spark applications in Python to perform cleansing, transformation, and enrichment of the data.
- Developed custom aggregate functions using Spark SQL and performed interactive querying.
- Designed a strategy for performing ETL on different events flowing in from Kinesis Firehose to S3.
- Created Kafka producer API to send JSON data into various Kafka topics.
- Have Experience with Databricks for processing and transforming massive quantities of data and exploring the data through machine learning models.
- Created aggregated tables on clickstream data using PySpark.
- Ingested Telecom data into HDFS from Teradata, FTP and upstream HDFS cluster using Sqoop, Bash scripts & Hadoop DistCp utility.
- Worked on Splunk for data logging, puppet and ansible to maintain physical Linux servers, and few machine learning applications to extract insights from data stored in HDFS.
- Showcased expertise in querying, visualization, and analytics tools to analyse and process complex data sets.
- Designed and implemented Hive UDFs for evaluation, filtering, loading, and storing of data.
- Involved in project Life Cycle - from analysis to production implementation, with emphasis on identifying the source and source data validation, developing logic and transformation as per the requirement and creating mappings and loading the data into different targets.
- Worked with GIT repository for all development and code maintenance.
- Engaged with offshore team in resolving the production issues.
Confidential
Hadoop Developer
Responsibilities:
- Implemented Partitioning, Dynamic Partitions, Buckets in HIVE.
- Developed the code for Importing and exporting data into HDFS and Hive using SQOOP.
- Worked extensively with SQOOP for importing and exporting the data from HDFS to DB2 Database systems and vice-versa loading data into HDFS.
- Written HIVE scripts as per the requirement.
- Designed and created managed/external tables in HIVE as per the requirement. Was involved in writing UDF’s in HIVE.
- Responsible for Turnover and promoting the code to QA, creating CR and CRQ for the release.
- Implementing POC to migrate map reduce jobs into Spark RDD transformations using Python.
- Had an opportunity to evaluate various tools to build several Proof of Concepts streaming / batch applications (Kafka) to bring in data from multiple data sources, transform and load data in to target systems and successfully implemented in Production.
- Very good exposure in building applications using Cloudera/Hortonworks Hadoop distributions.
- Created complex SQL queries and used JDBC connectivity to access the database.
- Built SQL queries to build the reports for presales and secondary sales estimations.
