We provide IT Staff Augmentation Services!

Sr. Data Engineer Resume

3.00/5 (Submit Your Rating)

DallaS

SUMMARY

  • More than 12+ years of experience on gathering System Requirements, Analyzing the requirements, Designing and developing systems.
  • Excellent Knowledge in understanding Big Data and Distributed file systems - HDFS, parallel processing - MapReduce framework and complete Hadoop ecosystem - Hive, Pig, Sqoop, Spark, Kafka, HBase, NoSQL, NiFi, Oozie and Flume.
  • 8+ years of experience on Big Data Technologies like Hadoop, Hive, Spark, Kafka, Sqoop, Oozie, Flume, HBase, Oozie, Flume and AWS.
  • Be able to design scalable, configurable, maintainable for complex business problems
  • Proficient in Business Analysis, Business Knowledge, Software Engineering Leadership, Architecture Knowledge and Technical Solution Design
  • Proactively, initiate, develop and maintain effective working relationships with team members. Coordinating with all team members, including 3rd party suppliers
  • Experience with agile/scrum methodologies to iterate quickly on product changes, developing user stories and working through backlog.
  • Hands on experience in writing HiveQL queries to do data cleansing and processing and experienced in hive performance optimization using Partitioning and Bucketing and Parallel Execution concepts.
  • Excellent understanding and knowledge on NOSQL databases like HBase, Cassandra and MongoDB.
  • Proficient in developing Sqoop scripts for the extractions of data from various RDBMS databases into HDFS.
  • Good working experience on different file formats like PARQUET, TEXTFILE, AVRO, ORC and different compression codecs GZIP, SNAPPY, LZO.
  • Strong ability to compile Java programming (Core Java) including OOPS concepts, Class, Method, Inheritance, Encapsulation, Loop, Exception handling etc.
  • Hands on experience in application development using Java, RDBMS (SQL), and Linux/Unix shell scripting.
  • Knowledge in implementing ETL/ELT processes with Hive and Spark.
  • Strong experience on multiple Hadoop distributions like Cloudera, Hortonworks and IBM BigInsights
  • Strong knowledge on batch and streaming data sources with structured and unstructured data
  • Strong knowledge on Data Warehouses, RDBMS and MPP database skills, including query optimization, and performance tuning
  • Leading large-scale data warehousing and analytics projects, including using AWS technologies - Redshift, EMR, S3, EC2, Lambda’s, Step functions, CloudFormation Template, SageMaker, Data-pipeline and other big data technologies.
  • Experience in working with AWS Code Pipeline and creating Cloud Formation JSON templates to create custom sized VPC & migrate a production infrastructure into an AWS utilizing CodeDeploy, CodeCommit, OpsWorks.
  • Hands-on experience with building AWS Lambda functions using Java and creating deployment ZIP API packages, handlers, monitoring the AWS Lambda Java functions through CloudWatch.
  • Hands on experience on Spark Framework with Spark core, Spark Streaming, Spark SQL for data processing by using Scala programing language.
  • Experienced Hadoop/Java developer and Spark/Scala having end to end experience in developing applications in Hadoop ecosystem.
  • Good knowledge of all phases of the Iterative Software Development Life Cycle (SDLC) and Strong independent team player.
  • Create and execute unit tests and perform basic application testing
  • Experience with Oozie Workflow Engine in running workflow jobs with actions that run Hadoop Spark and Hive jobs.
  • Experience working on Version control tools like SVN and GIT revision control systems such as GitHub and JIRA to track issues.
  • Cost optimization of the existing Data pipelines using Sparklens.
  • Fine tuning applications and systems for high performance and higher volume throughput.
  • Manages maintenance of applications and performs technical change requests scheduled according to Release Management processes
  • Monitor implementations to help ensure adherence to established standards
  • Develop best practices for developing and deploying Hadoop applications and assist the team to manage compliance to the standards
  • Execute change management activities supporting production deployment to Developers, Quality Control Analysts, and Environment Management personnel
  • Expertise on source code and version control management tools like Subversion and GIT and used Source code management client tools like Bitbucket, Git Bash, GitHub Desktop, GitLab and Git GUI.

TECHNICAL SKILLS

Big Data Ecosystemts: HDFS, Map-Reduce, Hive, Pig, Sqoop, NiFi, Flume &Zookeeper, Spark, Kafka, Flume

Language: Java, Scala, Python

Scripting: Shell Scripting and XML

RDBMS/Databases: Oracle10g and MS-SQL Server

Operating Systems: Windows2003, UNIX, Linux

Build Tools: SBT, Maven

Version Control Tools: GIT and BitBucket

PROFESSIONAL EXPERIENCE

Confidential, Dallas

Sr. Data Engineer

Responsibilities:

  • Performed Data ingestion, Batch Processing, Data Extraction, Transformation, Loading and Real Time Streaming using Hadoop Frameworks
  • The data is migrated from RDBMS data sources to Hive, Pig and HDFS by using Sqoop.
  • Exported the analyzed data to the relational databases using Sqoop for visualization and to generate reports by our BI team.
  • Involved in architecture design and implementation of best solutions with Hadoop/Big data based up on the requirements of the client.
  • Performance optimization and AWS cost optimization using Sparklens
  • Extensively worked on Hive tables, partitions and buckets for analyzing large volumes of data.
  • Used Flume to collect, aggregate, and store the web log data from different sources like web servers and network devices and pushed to HDFS.
  • Configure, monitor and automate Amazon Web Services as well as involved in deploying the content cloud platform on Amazon Web Services using EC2, S3 and EBS.
  • Hands-on experience with building AWS Lambda functions using Java and Scala and creating deployment ZIP API packages, handlers, monitoring the AWS Lambda functions through CloudWatch.
  • Integrated both framework andCloudFormation to automate AWS environment creation along with ability to deploy on AWS, using build scripts (AWS CLI) and automate solutions using Shell and Python.

Confidential, Philadelphia

Sr. BigData Developer

Responsibilities:

  • Worked on ingesting the source data into the Hadoop datalake (OLONA) from various databases by using Sqoop tool.
  • Reading, Processing and parsing CSV source data files through HQL script and ingesting to Hive and Impala tables.
  • Extensively worked on Hive, HBase, Impala tables, partitions and buckets for analyzing large volumes of data.
  • Scheduled the Hive queries daily by using oozie coordinator and by writing an oozie workflow.
  • I also worked on database testing and QA validation to make sure that the product is bug free.
  • Developed the application using Agile Methodology.
  • Proactively involved in ongoing maintenance, support and improvements in Hadoop cluster.

Confidential

AWS Developer

Responsibilities:

  • Design and develop the application using Agile/Scrum Methodologies.
  • Configure, monitor and automate Amazon Web Services as well as involved in deploying the content cloud platform on Amazon Web Services using EC2, S3 and EBS.
  • Hands-on experience with building AWS Lambda functions using Java and Scala and creating deployment ZIP API packages, handlers, monitoring the AWS Lambda functions through CloudWatch.
  • Created Terraform Scripts to Automate AWS services which include web servers, ELB, Cloud front Distribution, database, EC2 and database security groups, S3 bucket and application configuration, this Script creates stacks, single servers or joins web servers to stacks, AWS EKS and AWS Elastic Cache for Redis.

Confidential

Senior Data Engineer

Responsibilities:

  • Design and develop the application using Agile/Scrum Methodologies.
  • Involved in architecture design and implementation of best solutions with Hadoop/Big data based up on the requirements of the client.
  • Performed Data ingestion, Batch Processing, Data Extraction, Transformation, Loading and Real Time Streaming using Hadoop Frameworks.
  • Experience working with market data on Capital Market Project (Bloomberg).
  • Worked as an ETL/Big Data Lead with leading a team of 6 developers.
  • Interacted with multiple teams (Business Analyst, Project Management and Upstream development teams) and progressively tracking the issues and solving them.
  • Reading, Processing and parsing the source data files through Spark/Scala and ingesting to Hive tables.
  • Kafka streaming with Spark framework is implemented for ingesting and analyzing the huge volumes of Confidential data which is coming as 25TPS from the source systems.
  • Loading the data in to Hadoop ecosystem (Hive and HDFS)
  • Extensively worked on Hive tables partitions and buckets for analyzing large volumes of data.
  • Version controlling by using GIT Hub, Bitbucket tools and document maintenance by using JIRA, Confluence tools.
  • Worked in AWS Cloud IaaS stage with components VPC, ELB, Auto-Scaling, EBS, AMI, EMR, Redshift, Lambda, CloudFormation template, CloudFront, CloudTrail, CloudWatch and DynamoDB.
  • Reading, Processing and Parsing CSV source data files through Spark/Scala and ingesting to Hive tables.
  • The CSV source data is read through the Scala programming language from core (creating RDDs, Data Frame, Dataset, Scala Methods, Scala Classes and Objects, Pattern Matching, Working with Lists, Collections, Etc.)
  • Knowledge transition to the end users and junior developers to understand the hive queries and the business requirements.

Confidential, Edmonton

Senior Hadoop Developer

Responsibilities:

  • Design and develop the application using Agile Methodologies.
  • Loading the data into Hadoop ecosystem (Hive and Impala) by using Pentaho ETL tool (Spoon).
  • Extensively worked on Hive tables, Impala tables partitions and buckets for analyzing large volumes of data.
  • Version controlling by using GIT Hub, Source Tree tools and document maintenance by using JIRA, Confluence tools.
  • Jenkins tools are used for continuous integration services for software development and automated builds.
  • Used Apache parquet with Hive to make the advantages of compressed, efficient columnar data representation available to this project in Hadoop ecosystem.
  • Reading, Processing and parsing semi structured source data files through Spark/Scala and ingesting to Hive tables.
  • Knowledge transition to the end users and junior developers to understand the hive queries and the business requirements
  • Proactively involved in ongoing maintenance, support and improvements in Hadoop cluster.

Confidential

Senior Hadoop Developer

Responsibilities:

  • Worked on ingesting the source data into the Hadoop datalake from various databases by using Sqoop tool.
  • Reading, Processing and parsing CSV source data files through Spark/Scala and ingesting to Hive tables.
  • Extensively worked on Hive tables, partitions and buckets for analyzing large volumes of data.
  • Scheduled the Hive queries daily by using oozie coordinator and by writing an oozie workflow.
  • I also worked on database testing and QA validation to make sure that the product is bug free.
  • Knowledge transition to the end users and junior developers to understand the hive queries and the business requirements.
  • Developed the application using Agile Methodology.
  • Proactively involved in ongoing maintenance, support and improvements in Hadoop cluster.

Confidential

Responsibilities:

  • Collaborating with business users/product owners/developers to contribute to the analysis of functional requirements.
  • Design and develop the application using Agile Methodologies.
  • Involved in loading data from the system generated data sources to HDFS and experienced in writing multiple java-based Map-Reduce jobs for cleaning, processing the data.
  • Extensively worked on Hive tables, partitions and buckets for analyzing large volumes of data.
  • Used Apache Avro with Hive to make the advantages of compressed, efficient columnar data representation available to this project in Hadoop ecosystem.
  • Had experience with Source Code Repository systems (SVN) and used revision control systems such as Git.
  • Debugging Map Reduce jobs using job history logs and syslog for tasks.
  • Developed shell scripts for adding process dates to the source files.
  • Transferring the log incoming log files to the Parser (specially written Java Code to load in to HDFS, HBASE) by using Kafka message broker.
  • Design & Develop ETL workflow using Oozie for business requirements which includes automating the extraction of data from MySQL database into HDFS using Sqoop scripts.
  • Splunk tool is used to retrieve the data from the hadoop cluster.
  • Extensively worked on Apache Flume to collect the logs and error messages across the cluster.
  • Created a Spark POC to capture user click stream data and find what topics they are interested in. Scala language is used in Spark project.
  • Scheduling the Hive jobs using Oozie and Falcon process files.
  • Performed defect co-ordination with both Development & Testing Teams.
  • Performed data analytics in Hive and then exported this metrics back to Oracle Database using Sqoop.
  • Installed Hadoop single node and multi node clusters on Red Hat OS to store the data into HDFS for preforming various Hadoop jobs.
  • Jenkins tools are used for continuous integration services for software development and automated builds.
  • Conducting root cause analysis and resolve production problems and data issues.
  • Proactively involved in ongoing maintenance, support and improvements in Hadoop cluster.

Confidential

Hadoop developer

Responsibilities:

  • Proactively monitored systems and services, architecture design and implementation of hadoop deployment, configuration management, backup, and disaster recovery systems and procedures.
  • Involved in Analyzing system failures, identifying root causes, and recommended course of actions. Documented the systems processes and procedures for future references.
  • Used Flume to collect, aggregate, and store the web log data from different sources like web servers and network devices and pushed to HDFS.
  • Performed Java MapReduce programs on log data to transform into structured way to find user location, age group, spending time.
  • The data is migrated from RDBMS data sources to Hive, Pig and HDFS by using Sqoop.
  • Exported the analyzed data to the relational databases using Sqoop for visualization and to generate reports by our BI team.
  • Jira bug tracking tool is used for logging the bugs.
  • ETL (Informatica) tool was used in this project hence I had exposure towards Informatica.
  • Tableau reporting tool was used to generate the reports from the Hadoop cluster.
  • Integrated Oozie with the rest of the Hadoop stack supporting several types of Hadoop jobs out of the box (such as Java map-reduce, Streaming map-reduce, Pig, Hive, Sqoop and Distcp) as well as system specific jobs (such as Java programs and shell scripts).

We'd love your feedback!