We provide IT Staff Augmentation Services!

Data Engineer Resume

0/5 (Submit Your Rating)

Reston, VA

SUMMARY

  • BigData Developer, Data Engineer with experience on ETL tools Informatica, Teradata, Oracle, MySQL OBIEE, Hadoop stack including Hive, Sqoop, HDFS, PySpark with Python/Scala and experience in Design, Development, Implementation, Testing and maintenance of Enterprise Data Warehouse (EDW), Data Marts, Data Lake.
  • Over 11 years of IT experience in software development, big data management, data modeling, data integration, implementation and testing of enterprise class systems spanning big data frameworks, advanced analytics and Java/J2EE technologies.
  • Over 5+ years of hands on experience in Hadoop components & Map Reduce programming for parsing and populating tables for Terabytes of data.
  • Having 11 years of experience in all aspects of Software development including requirement analysis.
  • Strong technical/analytical skills with clear understanding of ETL architecture, design based on requirements
  • Implemented DataLake on Hadoop using Hive, Sqoop, Spark with Python which involved in migrating data from different sources including Oracle, MySql, Flat files, XML and JSON files to HDFS and Hive.
  • Supporting data scientist team by extracting, cleaning, preprocessing, creating data pipelines on DataLake.
  • Experience on extracting, transforming and loading large structured, semi - structured and unstructured data.
  • Worked on all phases of data warehouse development lifecycle, ETL design, implementation, optimization.
  • Migrating data using Sqoop, Spark from RDBMS to HDFS and vice-versa to Text, Parquet, Orc and Avro.
  • Experience in data cleansing and data mining, capturing data from existing RDBMS databases using Sqoop.
  • Familiar with data architecture including data ingestion pipeline design, Hadoop information architecture, data modeling and data mining, machine learning and advanced data processing.
  • Experienced in handling large datasets using Partitions, Spark in Memory capabilities, Broadcasts in Spark, Effective & efficient Joins, Transformations and other during ingestion process.
  • Proficient in data warehousing techniques and extensively working on complex mappings with SCD Type 1, Type 2, Type 3 and CDC (change data capture), Audit process, Error Handling Techniques for ETL process.
  • Worked on Planning, designing and developing multiple Data Warehouses and Data Marts projects.
  • Expertise in performance tuning, gap analysis, debugging, identifying bugs/defects in existing mappings/workflows by analyzing the data flow and evaluating transformations.
  • Experience in writing Pyspark programs, Sqoop scripts, developed Streamsets data pipelines to process data from multiple RDBMS systems to bigdata environment. Automated informatica/pyspark/Sqoop jobs using Control-M scheduler.
  • Experience in analyzing business requirements and translating to functional/technical design specifications.
  • Created Shell scripts to automate and schedule Informatica jobs, Teradata scripts based on requirements.
  • Good experience with Dimension modeling techniques like Star schema, Hybrid schemas and good knowledge on SDLC methodologies.

TECHNICAL SKILLS

  • Teradata
  • Oracle
  • SQL Server
  • MySQL
  • HDFS ETL and Query tools
  • Informatica
  • Hadoop
  • Hive
  • PySpark
  • Spark
  • Sqoop
  • SQL
  • Perl
  • Python
  • Bash
  • NiFi
  • StreamSets
  • Cloudera

PROFESSIONAL EXPERIENCE

Confidential, Reston, VA

Data Engineer

Responsibilities:

  • Migration RDBMS application to Hadoop platform.
  • Move all the ETL existing jobs to Hadoop.
  • Knowledge on NIFI.
  • Conversion of existing Teradata to big data platform using Hadoop, Pyspark, Sqoop, Hive, Impala, Unix shell scripting.
  • Validating, debugging and fixing issue related to model input dataset, custom features, model predicted scores.
  • Familiar with data architecture including data ingestion pipeline design, Hadoop information architecture, data modeling and data mining, machine learning and advanced data processing.
  • Strong experience in analyzing/understanding the business requirements and converting to technical requirements & implementing the code using Mapping documents.
  • Analyzing the source data, performing data profiling and extracting data from multiple RDBMS sources, cleansing, transforming and loading into data warehouse which is Oracle Exadata system.
  • Processing XML, Json, Delimited files using Informatica power center & transforming and loading into data warehouse/Data Mart.
  • Development/Enhancement of Informatica Mappings using various transformations for enhancements and improved reusability of the code using mapplets/ worklets and variables/parameters concepts.
  • Converting informatica ETL mappings to Pyspark programs, applying all the transformation logic in pyspark (Spark with python) and loading data into HDFS (Hadoop Distributed File System) in parquet file format.
  • Creating Hive tables & strong experience in writing Oracle SQL/PL-SQL queries, procedures and HQL Spark SQL.
  • Creating, refreshing hive tables and views in Impala in Cloudera for business users for the data analytics.
  • Involved from the scratch on bigdata environment setup for EDW projects, worked with admins on roles creation, project folders creation on edge node, as well as HDFS.

Confidential, Herndon, VA

Data Engineer

Responsibilities:

  • Processed data into HDFS by developing solutions.
  • Analysed the data using Map Reduce, Pig, Hive and produce summary results from Hadoop to downstream systems.
  • Worked extensively with HIVE DDLs and Hive Query language (HQLs).
  • Developed UDF functions and implemented it in HIVE Queries.
  • Implemented SQOOP for large dataset transfer betweenHadoopand RDBMs.
  • Created Spark Jobs to convert the periodic of XML data into a partition avro Data.
  • Used Sqoop widely in order to import data from various systems/sources (like MySQL) into HDFS.
  • Created components like Hive UDFs for missing functionality in HIVE for analytics.
  • Developing Scripts and Batch Job to schedule a bundle (group of coordinators) which consists of various business
  • Exported the analysed data from storage system and generated report for visualization purpose and handed to reporting team.
  • Involved in ETL, Data Integration and Migration.
  • Used different file formats like Text files, Sequence Files, Avro.
  • Cluster co-ordination services through Zookeeper.
  • Assisted in creating and maintaining Technical documentation to launching HADOOP Clusters and even for executing Hive queries and Pig Scripts.
  • Assisted in Cluster maintenance, cluster monitoring, adding and removing cluster nodes and Trouble shooting.
  • Helped in configuring of Hadoop cluster, Spark, HDFS, Developed multiple Map Reduce jobs in scala for data cleaning and pre-processing.

Confidential, Dallas, TX

Big Data/Hadoop developer

Responsibilities:

  • Working on cloudera Hadoop using, spark with python/scala, Kafka, Hive, MySql, Unix shell scripting.
  • Perform data cleaning transformation filtering, tagging, joining, parsing, and normalizing data sets throughout the end-to-end process.
  • Building data pipelines with data from event streams, NoSQL, APIs etc by leveraging Apache Spark, kafka, Scala, Hive on cloudera hadoop cluster .
  • Extracting structured data from MySql, semi-structured JSON data from Microservice API in batch and real-time streaming data and storing in hive tables and HDFS location
  • Establishing and following leading practices around secure code development and testing to ensure the platform is free of most common coding vulnerabilities.
  • Created Hive tables using Parquet, ORC files with LZO, Snappy compression to optimize IO, CPU and memory usage while processing.
  • Worked on performance optimization of data access by avoiding scanning large tables using Hive-based partitioning and bucketing, indexing tables, hive vectorized execution.
  • Experienced in handling large datasets using Partitions, Spark in Memory capabilities, Broadcasts in Spark, Effective & efficient Joins, Transformations and other during ingestion process.
  • Interaction with Admin for Kerberos Authentication and increase the size of memory for better performance.
  • Deployment the data to UAT using Jenkins build, ansible tower.
  • Used Spark-SQL to Load JSON data and create Schema RDD in PySpark and loaded it into Hive Tables and handled Structured data using SparkSQL.
  • Converted all SQL queries into Spark SQL using python for fast processing.

Confidential, Lansing MI

Software Developer

Responsibilities:

  • Involved in Analysis, designing and testing support.
  • Involved in coding for DAOs, Services and Controllers.
  • Used Rational Clear Case and Clear Quest tools for creating parent and child activities and for monitoring the tasks status as to whether the tasks have been successfully unit tested or not.
  • Used JavaScript for Client validations.
  • Used Junit for Unit testing the application.
  • Worked on HTML5, CSS3 JavaScript, AJAX, jQuery, React JS, Node Js, Bootstrap,
  • JSON, XML.
  • Extensively used oracle SQL and used spring data for mapping repository.
  • Work with users and business group to get the details regarding the issues like screenshots of the issue, detailed workflow to replicate the issue, user access/roles to the application and any other details that might help to resolve the issues.
  • Analyse/ Debug the code to find out the root cause of the issue and present the root cause to business and make the changes to the solution as per the business need to fix the issue.
  • Participate in knowledge transfer to ensure better grasp of the product and domain.
  • Monitor process and software changes that impact production support, communicate project information to the production support staff and raise production support issues to the project team.
  • Responsible for coaching and mentoring less experienced team members.

Environment: Java, JSP, Spring, Hibernate, JPA, HTML5, CSS3 JavaScript, AJAX, jQueryReact JS, Node Js, Bootstrap, JSON, XML, SQL Developer, JIRA, Apache Tomcat Application Server, Eclipse, Junit, soap UI, Atlassian’s tools.

We'd love your feedback!