Data Engineer Resume
Houston, TX
SUMMARY
- 7+ years of professional experience in IT in Analysis, Design, Development, Testing, Documentation, Deployment, Integration, and Maintenance of web based and Client/Server applications using Java and Big Data technologies.
- Hadoop Developer wif 5 years of working experience on designing and implementing complete end - to-end Big Data/Hadoop Infrastructure using Spark, HDFS, HIVE, Hbase, Sqoop, Kafka, Flume and Oozie.
- Good noledge of Spark and Hadoop Architecture and various components such as HDFS, YARN, Name Node, Data Node and MapReduce concepts
- Experience in analyzing data using Hive QL and custom MapReduce programs in Java
- Exploring wif theSparkfor improving the performance and optimization of the existing algorithms in Hadoop usingSparkContext,Spark Streaming, Spark-SQL, Data Frame, Pair RDD's,SparkYARN.
- Experience in importing and exporting data using Sqoop from Relational Database Systems to HDFS and vice-versa.
- Experience in working wif Azure, Cloudera and HortonWorks Hadoop Distributions.
- Experienced in working wif different file formats like Text, Sequence, Xml, Avro and Parquet.
- Configured Flume to extract the data from the web server output files to load into HDFS.
- Working noledge on NoSQL databases like HBase, Cassandra.
- Experienced in using Integrated Development environments.
- Migration from different databases (i.e. Oracle, DB2, MYSQL) to Hadoop.
- Prior experience working as Software Developer in Java/J2EE and related technologies.
- Experience in designing and coding web applications using Core Java and J2EE Technologies
- Excellent noledge in Java and SQL in application development and deployment.
- Experience wif Agile Methodology, Scrum Methodology and release management.
- Working experience of control version tools like SVN, Git.
- Extensive experience in Spark, Nifi and Hive.
- Experience wif Sqoop and Kafka.
- Good Experience wif -Oozie and Zookeeper
- Good Experience wif Azure, Cloudera and Hortonworks -Hadoop Distribution.
- Experience wif Oracle, MySQL, hbase and Cassandra.
- Extensive experience in JSPs, Servlets and JDBC.
TECHNICAL SKILLS
Big Data Technologies: Apache Spark, Apache Hadoop, YARN, Pig, Hive, Sqoop, Flume,Zookeeper, Kafka, Oozie
Cloud Technologies: Azure Kubernetes Services (AKS), Function App, KeyVault, SQL Server, Azure Synapse, Blob storage, Azure Tables, Azure Data Factory
BI tools: Tableau, Jmp
Languages: Java, Scala, python, Unix Shell Scripting
Operating System: Windows, UNIX, Linux
Web/App Servers: Apache Tomcat, Glass Fish, Websphere Application Server
IDE: Eclipse, NetBeans, IntelliJ
Web Services: RESTful
Build Tools: Maven, Ant, Sbt
DataBase: Relational: Oracle, PostgreSQL, MySQL, Teradata, SQL Server. NOSQL: HBase, Cassandra
Methodologies: Agile, Waterfall
Version Control: SVN, Git and IBM Rational Clear Case tool
PROFESSIONAL EXPERIENCE
Data Engineer
Confidential, Houston, TX
Responsibilities:
- Reading CSV files from Cygnet adapter as DataFrames by providing a custom defined schema.
- Developed Spark/Scala job that implements data quality rules which includes filtering duplicates, removing nulls and so on specified by the client.
- Loading filtered data to Phoenix/HBase by using Spark-Phoenix plugin.
- Implementing flattening of the data (converting vertical tables to horizontal tables) through pivot functionality.
- Developed the carry forwarding values if the current values are missing through Spark RDD’s and DataFrames.
- Compacting the HBase tables to has better storage capacity and read performance.
- Parallelization of Phoenix queries to load data from Phoenix through Spark.
- Develop Notebooks using Azure databricks and Spark
- Schedule the the jobs through Azure Data Factory by creating appropriate pipelines
- Loading data from Phoenix/HBase to Teradata by using Teradata Parallel Transfer TPT.
- Tuning the solution by performing appropriate caching of data and various solutions to has better performance.
Environment: (s): Azure Blob storage, Spark, Spark SQL, Scala, HDFS, Cloudera Data Platform, HBase, Phoenix, Zookeeper, Azure Data Factory, Azure tables, Azure Databricks, Teradata, Linux, Maven, ScalaTest (FunSuite), GIT, Intellij, Log4j
Spark / Hadoop Developer
Confidential, Atlanta, GA
Responsibilities:
- Developed Spark jobs that filter unnecessary records and find out unique records based on different criteria.
- Experienced in implementing POC's to migrate iterative map reduce programs into Spark transformations usingScala.
- Worked on Sequence files, RC files, Map side joins, bucketing, partitioning for Hive performance enhancement and storage improvement.
- Importing and exporting data into HDFS and Hive using Sqoop.
- Imported streaming data using Apache Kafka and Spark streaming into HDFS and designed Hive tables on top.
- Good experience in Hive partitioning, bucketing and perform different types of joins on Hive tables and implementing Hive SerDe’s like REGEX, JSON and Avro.
- Moving the data from Oracle, MS SQL Server in to HDFS using Sqoop and importing various formats of flat files in to HDFS.
- Utilized Hive partitioning, Bucketing and performed various kinds of joins on Hive tables.
- Involved in creating Hive external tables to perform ETL on data that is produced on daily basis.
- Improving the performance and optimization of existing algorithms in Hadoop using Spark context, Spark-SQL and Spark YARN using Scala.
- Implemented Hive Generic UDF's to implement business logic.
- Implemented test scripts to support test driven development and continuous integration.
- Develop and maintain several batch jobs to run automatically depending on business requirements.
Environment: Azure, Spark, HDFS, Hortonworks Hadoop Distribution, Hive, Sqoop, Zookeeper, Oozie, Spark Streaming, Kafka, Java, Linux, Maven, Oracle 11g/10g, Git, HTML, IntelliJ, UNIX Shell Scripting, Kerberos, Scala 2.11.8.
Hadoop Developer
Confidential, Houston, TX
Responsibilities:
- Worked on Hortonworks platform. Developed data pipeline using Flume and Sqoop to ingest customer behavioral data from traditional databases into HDFS for analysis.
- Ingested large volumes of data from Teradata to Hadoop using Sqoop.
- WritingTeradata sql queriesto join or performing any modifications in the table.
- Implemented a real time data flow to retrieve data form Pason data hub and store them in Impala and Kafka.
- Developed processors to retrieve list of different objects, from WITSML server.
- Developed a processor to get real time data by implementing pagination query.
- Developed a controller service to connect the Pason datahub.
- Implemented distributed map cache service in Nifi to maintain the state of different objects.
- Converted nested json to Flat json by shift operations through JoltTransformJson.
- Storing data in Impala through putSQL and by providing respective JDBC drivers.
- Sending real time log data to appropriate Kafka topics dynamically through publishKafka.
Environment: Hadoop 2.2.0, Map Reduce, Kafka, Yarn, Hive, Pig, Oozie, Sqoop, Flume, Oracle 11g, Core Java, Hortonworks, HDFS, Eclipse.
Java Developer
Confidential
Responsibilities:
- Involved in design, development and testing of the application.
- Extensively worked wif Spring MVC for developing J2EE Components.
- Developed servlets and JSPs wif Custom Tag Libraries for control of the business processes in the middle-tier and was involved in their integration.
- Involved in writing the test cases for the application using JUnit.
- Involved in creating various Data Access Objects for Addition, modification and deletion of records using various specification files.
- Expertise in creating databases, users, tables, triggers, macros, views, stored procedures, functions, joins and hash indexes inTeradatadatabase.
- Responsible for creating Restful Web services using JAX-RS.
- Developed required stored procedures and database functions using PL/SQL.
- Continuous Integration is done using Jenkins to continuously integrate code and to do the builds.
- Added logging and debugging capabilities using Log4j and using SVN.
- Interacted wif the client directly while capturing the requirements and project closure.
Environment: Java, JSP, HTML, Spring, JavaScript, CSS, Restful Web services, Eclipse, Hibernate, Teradata, SVN, Quality Center, LOG4j, XQuery Tomcat Server.
