We provide IT Staff Augmentation Services!

Sr.dataengineer Resume

4.00/5 (Submit Your Rating)

Richmond, VA

SUMMARY:

  • Professional Software developer with 8+ years of technical expertise in all phases of Software development cycle (SDLC), in various Industrial sectors expertizing in Big Data analyzing Frameworks and Java/J2EE technologies
  • 4+ years of industrial experience in Big Data analytics, Data manipulation, using Hadoop Eco system tools Map - Reduce, HDFS, Yarn/MRv2, Pig, Hive, HDFS, HBase, Spark, Kafka, Flume, Sqoop, Flume, Oozie, Avro, Sqoop, AWS, Spark integration with Cassandra, Avro, Solr and Zookeeper.
  • Hands on expertise in working and designing of Row keys & Schema Design with NOSQL databases like Mongo DB, HBase, Cassandra.
  • Extensively worked on Spark using Scala on cluster for computational (analytics), installed it on top of Hadoop performed advanced analytical application by making use of Spark with Hive and SQL/Oracle.
  • Excellent Programming skills at a higher level of abstraction using Scala, Java and Python.
  • Experience in using D-Streams, Accumulator, Broadcast variables, RDD caching for Spark Streaming.
  • Hands on experience in developing SPARK applications using Spark tools like RDD transformations, Spark core, Spark MLlib, Spark Streaming and Spark SQL.
  • Strong experience and knowledge of real time data analytics using Spark Streaming, Kafka and Flume.
  • Working knowledge of Amazon’s Elastic Cloud Compute(EC2) infrastructure for computational tasks and Simple Storage Service (S3) as Storage mechanism.
  • Running of Apache Hadoop, CDH and Map-R distributions, Elastic MapReduce(EMR) on (EC2).
  • Expertise in developing Pig Latin scripts and Hive Query Language.
  • Developed Customized UDFs and UDAF’s in java to extend HIVE and Pig core functionality.
  • Created Hive tables to store structured data into HDFS and processed it using HiveQL.
  • Experience in validating and cleansing the data using Pig statements and hands-on experience in developing Pig MACROS.
  • Excellent Programming skills at a higher level of abstraction using Scala, Java and Python.
  • Working knowledge in installing and maintaining Cassandra by configuring the cassandra.yaml file as per the business requirement and performed reads/writes using Java JDBC connectivity.
  • Written multiple MapReduce Jobs using Java API, Pig and Hive for data extraction, transformation and aggregation from multiple file formats including Parquet, Avro, XML, JSON, CSV, ORCFILE and other compressed file formats Codecs like gZip, Snappy, Lzo.
  • Good experience in optimizing MapReduce algorithms using Mappers, Reducers, combiners and partitioner’s to deliver the best results for the large datasets.
  • Good knowledge on build tools like Maven, Log4j and Ant.
  • Hands on experience in using various Hadoop distributions (Cloudera (CDH 4/CDH 5), Hortonworks, Map-R, IBM Big insights, Apache and Amazon EMR Hadoop distributions.
  • Experienced in writing Ad Hoc queries using Cloudera Impala, also used Impala analytical functions.
  • In depth understanding/knowledge of Hadoop Architecture and various components such as HDFS, MapReduce Programming Paradigm and YARN architecture.
  • Proficient in developing, deploying and managing the Solr from development to production.
  • Used various Project Management services like JIRA for tracking issues, GitHub for various code reviews and Worked on various version control tools like CVS, GIT, SVN.
  • Hands-on knowledge in Core Java concepts like Exceptions, Collections, Data-structures, I/O. Multi-threading, Serialization and deserialization of streaming applications.
  • Experience in Software Design, Development and Implementation of Client/Server Web based Applications using JSTL, jQuery, JavaScript, Java Beans, JDBC, Struts, PL/SQL, SQL, HTML, CSS, PHP, XML, AJAX and had a bird’s eye view on React JavaScript Library.
  • Experience in maintaining an Apache Tomcat MYSQL, LDAP, Web service environment.
  • Designed ETL workflows on Tableau, Deployed data from various sources to HDFS.
  • Done Clustering, regression and Classification using Machine learning libraries Mahout, MLlib(Spark).
  • Good experience with use-case development, with Software methodologies like Agile and Waterfall.
  • Proven ability to manage all stages of project development Strong Problem Solving and Analytical skills and abilities to make Balanced & Independent Decisions.

TECHNICAL SKILLS:

Big Data Ecosystem: HDFS, MapReduce, Pig, Hive, Spark, YARN, Kafka, Flume, Sqoop, Solr, Impala, Oozie, ZooKeeper, Spark, Ambari, Mahout, MongoDB, Cassandra, Avro, Parquet and Snappy.

Hadoop Distributions: Cloudera (CDH3, CDH4, and CDH5), Hortonworks, MapR and Apache

Languages: Java, Python, Scala, SQL, JavaScript and C/C++

No SQL Databases: Cassandra, MongoDB, HBase and Amazon Dynamo dB.

Java Technologies: JSE, Servlets, JavaBeans, JSP, JDBC, JNDI, AJAX, EJB and struts

Web Design Tools: HTML, DHTML, AJAX, JavaScript, JQuery and JSON

Development / Build Tools: Eclipse, Jenkins, Git, Ant, Maven, IntelliJ, JUNIT and log4J.

App/Web servers: WebSphere, WebLogic, JBoss and Tomcat

DB Languages: MySQL, PL/SQL, PostgreSQL and Oracle

RDBMS: Oracle 10g,11i, MS SQL Server, MySQL and DB2

Operating systems: UNIX, Red Hat LINUX, Mac os and Windows Variants

Testing: Junit

ETL Tools: Talend

PROFESSIONAL EXPERIENCE:

Confidential, Richmond, VA

Sr.DataEngineer

Responsibilities:

  • Responsible for building scalable distributed data solutions using Hadoop.
  • Designing and developing Java codes for various job requirements to perform MapReduce jobs.
  • Performed advanced procedures like text analytics and processing, using the in-memory computing capabilities of Spark.
  • Created Lambda functions in AWS using Python scripts and created SNS triggers for the Lambda.
  • Experienced in handling large datasets using Partitions, Spark in Memory capabilities, Broadcasts in Spark, Transformations and other during ingestion process.
  • Worked with Spark to create structured data from the pool of unstructured data received in HDFS.
  • Developed Spark jobs to create data frames from source system, process and analyze the data in Data Frames based on business requirements.
  • Involved in continuous monitoring and managing the Hadoop cluster using Cloudera Manager.
  • Created test cases for each module in KYC Pipeline to test data quality and functionality.
  • Worked on interactive shell scripting for scheduling various data cleansing and data loading process into PostgreSQL.
  • Experienced in importing data from Hive tables and created data frames to join the customer data and storing into PostgreSQL database.
  • Used complex Nested Json Data for testing and development purpose.
  • Working experience in AWS (Amazon S3, EMR, EC2, LAMBDA, SNS, CLOUDWATCH).
  • Worked on live 20 node EMR cluster to run spark jobs to process data in Amazon S3 buckets.
  • Monitoring and Debugging Spark jobs which are running on a spark cluster using Cloudera Manager.

Environment: Java, Scala, Python, Hadoop, Apache Spark, MapReduce, Amazon Web Services, CDH 5.9, Cloudera Manager, Control M Scheduler, PostgreSQL, Shell Scripting, Agile Methodology, JIRA, Git, Tableau.

Client: Nielsen, Chicago, IL

Role: Sr.Spark/Python Developer

Responsibilities:

  • Ability to excel and succeed in diverse environments and projects with strong determination, dedication and inclination.
  • Involved in HBase table design and Hive queries to populate Hive tables for analysis. Involved in architecture build up for the data flow and data storage.
  • Experienced with deployments, maintenance and troubleshooting applications on Microsoft Azure Cloud infrastructure.
  • Develop MapReduce jobs to convert data files into Parquet file format.
  • Execute Hive queries on Parquet tables stored in Hive to perform data analysis to meet the business requirements.
  • Experience in performing SQL and hive operations using Spark SQL.
  • Used Bitbucket as a repository for storing the code and integrated with Jupyter Notebook for integration purpose.
  • Job Developed Apache Spark and Implemented Apache Spark data processing project to handle data from various RDBMS and Streaming sources.
  • Experience in developing various Spark Streaming Jobs and Scala.
  • Developing spark code to apply various transformations and actions for faster data processing.
  • Used Spark processing to get data into in-memory, implemented RDD transformations, and performed actions.
  • PIG and Hive scripts were created for data ingestion from relational databases to compare with historical data.
  • Experienced in migrating HiveQL into Impala to minimize query response time.
  • Knowledge on handling Hive queries using Spark SQL that integrates with Spark environment.
  • Extensive use of Parquet format and snappy codec compressions to save Storage and improve performance in HDFS.
  • Worked with sparksql to query data from Hive tables.
  • Extensively used Sqoop to extract data from legacy systems(oracle) and load the data into HDFS
  • Worked on migrating spark scripts, oozie workflows, hive .
  • Worked with different File Formats like textfile, Parquet for HIVE querying and processing based on business logic.
  • Used JIRA for creating the user stories and creating branches in the bitbucket repositories based on the story.
  • Knowledge on creating various repositories and version control using GIT.
  • Involved in story-driven agile development methodology and actively participated in daily scrum meetings.

Environment:: GIT, Azure, Map Reduce, HDFS, SQL, Hive, Impala, Spark, Oozie, Linux, Airflow, Kubernetes, Jira, Bitbucket.

Confidential, Columbus, OH

Sr. Spark/Scala Developer

Responsibilities:

  • Developed Spark Applications by using Scala and Implemented Apache Spark data processing project to handle data from various RDBMS and Streaming sources.
  • Worked with the Spark for improving performance and optimization of the existing algorithms in Hadoop using Spark Context, Spark-SQL, Spark MLlib, Data Frame, Pair RDD's, Spark YARN.
  • Used Spark Streaming APIs to perform transformations and actions on the fly for building common learner data model which gets the data from Kafka in Near real time and persist it to Cassandra.
  • Developed Kafka consumer's API in Scala for consuming data from Kafka topics.
  • Consumed XML messages using Kafka and processed the xml file using Spark Streaming to capture UI updates.
  • Developed Preprocessing job using Spark Data frames to flatten Json documents to flat file.
  • Load D-Stream data into Spark RDD and do in memory data Computation to generate Output response.
  • Experienced in writing live Real-time Processing and core jobs using Spark Streaming with Kafka as a data pipe-line system.
  • Used Apache Nifi for ingestion of data from the Messages Queue's
  • Migrated an existing on-premises application to AWS. Used AWS services like EC2 and S3 for small data sets processing and storage, Experienced in Maintaining the Hadoop cluster on AWS EMR.
  • Imported data from AWS S3 into Spark RDD, Performed transformations and actions on RDD's.
  • Good understanding of Cassandra architecture, replication strategy, gossip, snitch etc.
  • Designed TABLES in Cassandra and Ingested data from RDBMS, performed data transformations, and then exported the transformed data to Cassandra as per the business requirement.
  • Used the Spark DataStax Cassandra Connector to load data to and from Cassandra.
  • Experienced in Creating data-models for Client’s transactional logs, analyzed the data from Casandra tables for quick searching, sorting and grouping using the Cassandra Query Language(CQL).
  • Tested the cluster Performance using Cassandra-stress tool to measure and improve the Read/Writes.
  • Used Hive QL to analyze the partitioned and bucketed data, Executed Hive queries on Parquet tables stored in Hive to perform data analysis to meet the business specification logic.
  • Used Apache Kafka to aggregate web log data from multiple servers and make them available in Downstream systems for Data analysis and engineering type of roles.
  • Experience in using Avro, Parquet, RCFile and JSON file formats, developed UDFs in Hive and Pig.
  • Developed Custom Pig UDFs in Java and used UDFs from PiggyBank for sorting and preparing the data.
  • Developed Custom Loaders and Storage Classes in PIG to work on several data formats like JSON, XML, CSV and generated Bags for processing using pig etc.
  • Developed Sqoop and Kafka Jobs to load data from RDBMS, External Systems into HDFS and HIVE.
  • Developed Oozie coordinators to schedule Pig and Hive scripts to create Data pipelines.
  • Used Impala where ever possible to achieve faster results compared to Hive during Data Analysis.
  • Written several Map reduce Jobs using Java API, also Used Jenkins for Continuous integration.
  • Setting up and worked on Kerberos authentication principals to establish secure network communication on cluster and testing of HDFS, Hive, Pig and MapReduce to access cluster for new users.
  • Continuous monitoring and managing the Hadoop cluster through Cloudera Manager.

Environment: Spark, Spark-Streaming, Spark SQL, AWS EMR, MapR, HDFS, Impala, NiFi, Hive, Pig, Apache Kafka, Sqoop, Java (JDK SE 6, 7), Scala, Shell scripting, Linux, MySQL Oracle Enterprise DB, SOLR, Jenkins, Eclipse, Oracle, Git, Oozie, MySQL, Soap, Cassandra and Agile Methodologies.

Confidential, Houston, TX

Hadoop/Spark Developer

Responsibilities:

  • Worked on migrating MapReduce programs into Spark transformations using Spark and Scala, initially done using python (PySpark).
  • Developed Spark jobs using Scala on top of Yarn/MRv2 for interactive and Batch Analysis.
  • Experienced in querying data using SparkSQL on top of Spark engine for faster data sets processing.
  • Worked on implementing Spark Framework a Java based Web Framework.
  • Worked and learned a great deal from AWS cloud services like EC2, S3, EBS, RDS and VPC.
  • Implemented Elastic search on Hive data warehouse platform.
  • Worked with ELASTIC MAPREDUCE and setup Hadoop environment in AWS EC2 Instances.
  • Responsible for migrating the code base from Cloudera Platform to Amazon EMR and evaluated
  • Amazon eco systems components like RedShift, Dynamo DB.
  • Used Amazon Redshift for data warehouse and to generate backend reports.
  • Written java code to format XML documents, uploaded them to Solr server for indexing.
  • Optimized Hive QL Scripts by using execution engine like Tez.
  • Worked on Ad hoc queries, Indexing, Replication, Load balancing, Aggregation in MongoDB.
  • Processed the Web server logs by developing Multi-hop flume agents by using Avro Sink and loaded into MongoDB for further analysis, also extracted files from MongoDB through Flume and processed.
  • Used Scala and Python to convert Hive/SQL queries into RDD transformations in Apache Spark.
  • Experience in developing, support and maintenance for the ETL (Extract, Transform and Load) processes using Talend Integration Suite.
  • Experienced in Hadoop Big Data Integration with ETL on performing data extract, loading and transformation process for ERP data.
  • Experienced in analyzing, designing and developing ETL strategies and processes, writing ETL specifications.
  • Analyzing the source data to know the quality of data by using Talend Data Quality.
  • Expert knowledge on MongoDB NoSQL data modeling, tuning, disaster recovery backup used it for distributed storage and processing using CRUD.
  • Extracted and restructured the data into MongoDB using import and export command line utility tool.
  • Experience in setting up Fan-out workflow in flume to design v shaped architecture to take data from many sources and ingest into single sink.
  • Experience in creating tables, dropping and altered at run time without blocking updates and queries using HBase and Hive.
  • Experience in working with different join patterns and implemented both Map and Reduce Side Joins.
  • Wrote Flume configuration files for importing streaming log data into HBase with Flume.
  • Imported several transactional logs from web servers with Flume to ingest the data into HDFS. Using Flume and Spool directory for loading the data from local system(LFS) to HDFS.
  • Installed and configured pig, written Pig Latin scripts to convert the data from Text file to Avro format.
  • Developed Spark scripts by using Scala IDE as per the business requirement.
  • Created Partitioned Hive tables and worked on them using HiveQL.
  • Loading Data into HBase using Bulk Load and Non-bulk load.
  • Worked on continuous Integration tools Jenkins and automated jar files at end of day.
  • Worked with Tableau and Integrated Hive, Tableau Desktop reports and published to Tableau Server.
  • Developed MapReduce programs in Java for parsing the raw data and populating staging Tables.
  • Experience in setting up the whole app stack, setup and debug logstash to send Apache logs to AWS Elastic search .
  • Worked on querying tools like Hue, Zeeplin, Presto and Impala.
  • Used Zookeeper to coordinate the servers in clusters and to maintain the data consistency.
  • Experienced knowledge over designing Restful services using java based API’s like JERSEY.
  • Experience in importing and exporting data using Sqoop from HDFS to Relational Database Systems and vice-versa and have experience in using apache NiFi to copy the data from local file system to HDFS.
  • Used OOZIE Operational Services for batch processing and scheduling workflows dynamically.
  • Supported in setting up QA environment and updating configurations for implementing scripts with Pig, Hive and Sqoop

Environment: HDP 2.3, Hadoop, HDFS, Hive, Map Reduce, AWS Ec2, SOLR, Redshift, Impala, MySQL, Oracle, Sqoop, Flume, Spark, SQL Talend, Python, PySpark, Yarn, Pig, Oozie, Linux-Ubuntu, Scala, Ab Initio, Tableau, Maven, Jenkins, Java (JDK 1.6), Cloudera, JUnit, agile methodologies

Confidential, Chesterfield, MO

Big Data Hadoop Developer

Responsibilities:

  • Analyzing and writing Hadoop Map reduce jobs using Java API, Pig and Hive.
  • Exported data using Sqoop from HDFS to Teradata on regular basis.
  • Write scripts to automate application deployments and configurations. Monitoring YARN applications.
  • Wrote map reduce programs to clean and pre-process the data coming from different sources.
  • Implemented various output formats like Sequence file and parquet format in Map reduce programs. Also, implemented multiple output formats in the same program to match the use cases.
  • Using Pig to apply transformations, cleaning and deduplication of data from raw data sources.
  • Installation of Oozie workflow to run multiple Hive.
  • Implemented test scripts to support test driven development and continuous integration.
  • Converted text files into Avro then to parquet format for the file to be used with other Hadoop eco system tools.
  • Experienced on loading and transforming of large sets of structured, semi structured and unstructured data.
  • Exported the analysed data to HBase using Sqoop and to generate reports for the BI team.
  • Worked with TIBCO Spotfire Statistical Services
  • Analysed large amounts of data sets to determine optimal way to aggregate and report on it.
  • Participate in requirement gathering and analysis phase of the project in documenting the business requirements by conducting workshops/meetings with various business users.

Environment: Hadoop 1.0.4, Python, MapReduce, HDFS, Hive 0.10, Pig, Hue, Spark, Tibco, Kafka, Oozie, Core Java, Eclipse, Hbase, Flume, Cloudera Manager, Greenplum DB, IDMS, VSAM, SQL*PLUS, Toad, Putty, Windows NT, UNIX Shell Scripting, Linux 5, Pentaho Big data, YARN, HawQ, SpringXD, Eclipse, Java SDK 1.6 .

Confidential

Java / Web Developer

Responsibilities:

  • Implemented the project according to the Software Development Life Cycle (SDLC).
  • Analysing and Preparing the requirement Analysis Document.
  • Involved in developing Web Services using SOAP for sending and getting data from external interface.
  • Involved in requirement gathering, requirement analysis, defining scope, and design.
  • Worked with various J2EE components like Servlets, JSPs, JNDI, JDBC using Web Logic Application server.
  • Involved in developing and coding the Interfaces and classes required for the application and created appropriate relationships between the system classes and the interfaces provided.
  • Assisting project managers with drafting use case scenarios during the planning stages.
  • Developing the Use Cases, Class Diagrams and Sequence Diagrams.
  • Used Java Script for client-side Validation.
  • Used HTML, CSS, JavaScript for create web pages.
  • Involved in Database design and developing SQL Queries, stored procedures on MySQL.

Environment:: Java, J2EE, JDBC, HTML, CSS, JavaScript, Servlets, JSP, JDBC, Oracle, Eclipse, Web Logic, MySQL.

Confidential

Java Developer

Responsibilities:

  • Involved in Requirements Analysis, and design an Object-oriented domain model.
  • Involvement in the detailed Documentation, written functional specifications of the module.
  • Involved in development of Application with Java and J2EE technologies.
  • Develop and maintain elaborate services based architecture utilizing open source technologies like Hibernate, ORM and Spring Framework.
  • Developed server-side services using Java multithreading, Struts MVC, Java, EJB, Spring, Web Services (SOAP, WSDL, AXIS).
  • Responsible for developing DAO layer using Spring MVC and configuration XML’s for Hibernate and to also manage CRUD operations (insert, update, and delete).
  • Designing, Development and Implementation of JSPs in Presentation layer for Submission, Application, reference implementation.
  • Development of JavaScript for client end data entry validations and Front-End Validation.
  • Deployed Web, presentation and business components on Apache Tomcat Application Server.
  • Developed PL/SQL procedures for different use case scenarios
  • Involvement in post-production support, Testing and used JUNIT for unit testing of the module.

Environment: Java/J2EE, JSP, XML, Spring Framework, Hibernate, Eclipse(IDE), Java Script, Ant, SQL, PL/SQL, Oracle, Windows, UNIX, Soap, Jasper reports.

We'd love your feedback!