We provide IT Staff Augmentation Services!

Senior Big Data Engineer Resume

4.00/5 (Submit Your Rating)

Desmoines, IA

SUMMARY

  • Having 8+ years of experience asBig Data Engineer /Data EngineerandData Analystincluding designing, developing and implementation ofdata modelsfor enterprise - level applications and systems.
  • Experienced working with various Hadoop Distributions (Cloudera, Hortonworks, Map R, Amazon EMR) to fully implement and leverage new Hadoop features.
  • Experience in developingcustomUDFsfor Pig and Hive to in corporate methods and functionality of Python/Java intoPig LatinandHQL(HiveQL) and Used UDFs from Piggybank UDF Repository.
  • Experience on Migrating SQL database to Azure Data Lake, Azure data lake Analytics, Azure SQL Database, Data Bricks and Azure SQL Data warehouse and Controlling and granting database access and Migrating On premise databases to Azure Data lake store using Azure Data factory
  • Deep knowledge of troubleshooting and tuning Spark applications and Hive scripts to achieve optimal performance
  • Strong experience with ETL and/or orchestration tools (e.g. Talend, Oozie, Airflow)
  • Experience setting up AWS Data Platform - AWS CloudFormation, Development End Points, AWS Glue, EMR and Jupyter/Sagemaker Notebooks, Redshift, S3, and EC2 instances
  • Experienced in using Agile methodologies including extreme programming, SCRUM and Test-Driven Development (TDD)
  • Experience in developing Spark Applications using Spark RDD, Spark - SQL and Data frame APIs.
  • Worked with real-time data processing and streaming techniques using Spark streaming and Kafka.
  • Experience in moving data into and out of the HDFS and Relational Database Systems (RDBMS) using Apache Sqoop.
  • Good understanding of Partitions, bucketing concepts in Hive and designed both Managed and External tables in Hive to optimize performance.
  • Strong experience in core Java,Scala, SQL, PL/SQL and Restful web services.
  • Strong experience using HDFS, MapReduce, Hive, Spark, Sqoop, Oozie, and HBase.
  • Deep knowledge of troubleshooting and tuning Spark applications and Hive scripts to achieve optimal performance.
  • Expertise in working with HIVE data warehouse infrastructure-creating tables, data distribution by implementing Partitioning and Bucketing, developing and tuning the HQL queries.
  • Replaced existing MR jobs and Hive scripts with Spark SQL & Spark data transformations for efficient data processing.
  • Database design, modeling, migration and development experience in using stored procedures, triggers, cursor, constraints and functions. Used My SQL, MS SQL Server, DB2, and Oracle
  • Experience working with NoSQL database technologies, including MongoDB, Cassandra and HBase.
  • Experience with Software development tools such as JIRA, Play, GIT.
  • Used Informatica Power Center for (ETL) extraction, transformation and loading data from heterogeneous source systems into target database
  • Experience developing Kafka producers and Kafka Consumers for streaming millions of events per second on streaming data
  • Good understanding of the Data modelling (Dimensional & Relational) concepts like Star-Schema Modelling, a Schema Modelling, Fact and Dimension tables.
  • Experience in manipulating/analysing large datasets and finding patterns and insights within structured and unstructured data.
  • Strong understanding of Java Virtual Machines and multi-threading process.
  • Experience in writing complex SQL queries, creating reports and dashboards.
  • Proficient in using Unix based Command Line Interface.

TECHNICAL SKILLS

Big Data/ Hadoop Ecosystem: HDFS, Map Reduce YARN, Hive, Pig, Hbase, Kafka, Impala, Zookeeper, Sqoop, Oozie, DataStax & Apache Cassandra, Drill, Flume, Spark, Solr and Avro

Web Technologies: HTML, XML, JDBC, JSP, JavaScript, AJAX

RDBMS: Oracle 12c, MySQL, SQL server, Teradata

No SQL: Hbase, Cassandra, MongoDB

Web/Application servers: Tomcat, LDAP

Methodologies: Agile, UML, Design Patterns (Core Java and J2EE)

Cloud Environment: AWS, MS Azure

Development Tools: Microsoft SQL Studio, IntelliJ, Azure Databricks, Eclipse, NetBeans

Programming Languages: Scala, Python, SQL, Java, PL/SQL, Linux shell scripts.

Tools: Used: Eclipse, Putty, Cygwin, MS Office

BI Tools: Platfora, Tableau, Pentaho

PROFESSIONAL EXPERIENCE

Confidential, DesMoines, IA

Senior Big Data Engineer

Responsibilities:

  • Work in a fast-paced agile development environment to quickly analyze, develop, and test potential use cases for the business.
  • Experience in data processing like collecting, aggregating, moving the data using Apache Kafka.
  • Used Kafka to load data into HDFS and move data back to S3 after data processing
  • Developed Automation Regressing Scripts for validation of ETL process between multiple databases like AWS Redshift, Oracle, MongoDB, T-SQL, and SQL Server usingPython.
  • Involved as primary on-site ETL Developer during the analysis, planning, design, development, and implementation stages of projects using IBM Web Sphere software (Quality Stage v9.1, Web Service, Information Analyzer, Profile Stage)
  • Worked with Hadoop infrastructure to storage data in HDFS storage and use Spark / HIVE SQL to migrate underlying SQL codebase in AWS.
  • Worked with Hadoop ecosystem and Implemented Spark using Scala and utilized Dataframes and Spark SQL API for faster processing of data.
  • Developing Spark programs with Python, and applied principles of functional programming to process the complex structured data sets.
  • Converting Hive/SQL queries into Spark transformations using Spark RDDs and Pyspark
  • Analyzing SQL scripts and designed the solution to implement using PySpark
  • Export tables from Teradata to HDFS using Sqoop and build tables in Hive.
  • Loaded and transformed large sets of structured, semi structured and unstructured data usingHadoop/Big Data concepts.
  • Use SparkSQL to load JSON data and create Schema RDD and loaded it into Hive Tables and handled structured data using SparkSQL.
  • Generate metadata, create Talend ETL jobs, mappings to load data warehouse, data lake.
  • Designed and Developed Real Time Stream Processing Application using Spark, Kafka, Scala and Hive to perform Streaming ETL and apply Machine Learning.
  • Evaluated big data technologies and prototype solutions to improve our data processing architecture. Data modeling, development and administration of relational and NoSQL databases (Big Query, Elastic Search)
  • Utilized Spark, Scala, Hadoop, HBase, Cassandra, MongoDB, Kafka, Spark Streaming, a broad variety of machine learning methods including classifications, regressions, dimensionally reduction etc.
  • Collected data using Spark Streaming from AWS S3 bucket in near-real-time and performs necessary Transformations and Aggregation on the fly to build the common learner data model and persists the data in HDFS.
  • Involved in Relational and Dimensional Data modeling for creating Logical and Physical Design of Database and ER Diagrams with all related entities and relationship with each entity based on the rules provided by the business manager using ERWIN r9.6.
  • Developed Spark code using Scala and Spark-SQL/Streaming for faster processing of data.
  • Develop RDD's/Data Frames in Spark using and apply several transformation logics to load data from Hadoop Data Lakes.
  • The individual will be responsible for design and development of High-performance data architectures which support data warehousing, real-time ETL and batch big-data processing.
  • Filtering and cleaning data using Scala code and SQL Queries
  • Prepared Data Mapping Documents and Design the ETL jobs based on the DMD with required Tables in the Dev Environment.
  • Implemented Installation and configuration of multi-node cluster on Cloud using Amazon Web Services (AWS) onEC2.
  • Designed and developed architecture for data services ecosystem spanning Relational, NoSQL, and Big data technologies. Extracted Mega Data from Amazon Redshift, AWS, and Elastic Search engine using SQL Queries to create reports.
  • Used Talend for Big Data Integration using Spark and Hadoop.
  • UsedKafkaandKafka brokers, initiated the spark context and processed live streaming information with RDD and Used Kafka to load data into HDFS and NoSQL databases.
  • UsedZookeeperto store offsets of messages consumed for a specific topic and partition by a specific Consumer Group in Kafka.
  • UsedKafkafunctionalities like distribution, partition, replicated commit log service for messaging systems by maintaining feeds and Created applications using Kafka, which monitors consumer lag withinApache Kafkaclusters.
  • Involved with writing scripts in Oracle, SQL Server and Netezza databases to extract data for reporting and analysis and Worked in importing and cleansing of data from various sources like DB2, Oracle, flat files onto SQL Server with high volume data
  • Worked inAWSenvironment for development and deployment of custom Hadoop applications.
  • Strong experience in working withELASTIC MAPREDUCE(EMR) and setting up environments on AmazonAWSEC2 instances.
  • Worked on SQL Server concepts SSIS (SQL Server Integration Services), SSAS (Analysis Services) and SSRS (Reporting Services). Using Informatica & SSIS, SPSS, SAS to extract transform & load source data from transaction systems.
  • Developed reusable objects like PL/SQL program units and libraries, database procedures and functions, database triggers to be used by the team and satisfying the business rules.
  • Experience with Data Analytics, Data Reporting, Ad-hoc Reporting, Graphs, Scales, PivotTables and OLTP reporting.

ENVIRONMENT: Hadoop, Spark, Scala, Hbase, Hive, UNIX, Erwin, TOAD, MS SQL Server database, XML files, AWS, Cassandra, MongoDB, Kafka, IBM Info Sphere Data Stage, PL/SQL, Oracle 12c, Flat files, Autosys, MS Access database.

Confidential, Dania Beach, FL

Big Data Engineer

Responsibilities:

  • Worked with SCRUM team in delivering agreed user stories on time for every Sprint.
  • Worked on analyzing and resolving the production job failures in several scenarios.
  • Performed Data Preparation by using Pig Latin to get the right data format needed.
  • Participated in Data Acquisition with Data Engineer team to extract clinical and imaging data from several data sources like flat file and other databases.
  • Created Session Beans and controller Servlets for handling HTTP requests from Talend
  • Recreating existing application logic and functionality in the Azure Data Lake, Data Factory, SQL Database and SQL data warehouse environment.
  • Build machine learning models to showcase Big data capabilities using Pyspark and MLlib.
  • Enhancing Data Ingestion Framework by creating more robust and secure data pipelines.
  • Implemented data streaming capability using Kafka and Talend for multiple data sources.
  • Worked with multiple storage formats (Avro, Parquet) and databases (Hive, Impala, Kudu).
  • Processed the image data through the Hadoop distributed system by using Map and Reduce then stored into HDFS.
  • Working knowledge of cluster security components like Kerberos, Sentry, SSL/TLS etc.
  • Involved in the development of agile, iterative, and proven data modeling patterns that provide flexibility.
  • Knowledge on implementing the JILs to automate the jobs in production cluster.
  • UsedSCALAto storestreaming datato HDFS and to implementSparkfor faster processing of data.
  • Developed the Apache Storm, Kafka, and HDFS integration project to do a real time data analyses.
  • Implemented UNIX scripts to define the use case workflow and to process the data files and automate the jobs.
  • Used windows Azure SQL reporting services to create reports with tables, charts and maps
  • Populated HDFS and PostgreSQL with huge amounts of data using Apache Kafka.
  • Troubleshoot user’s analyses bugs (JIRA and IRIS Ticket).
  • Installed and configured ApacheHadoopto test the maintenance of log files in Hadoop cluster.
  • Installed and configuredHive, Pig, Sqoop, FlumeandOozieon the Hadoop cluster.
  • InstalledOozie workflowengine to run multiple Hive and Pig Jobs.
  • Setup and benchmarked Hadoop/Hbase clusters for internal use.
  • Architect & implement medium to large scale BI solutions on Azure using Azure Data Platform services (Azure Data Lake, Data Factory, Data Lake Analytics, Stream Analytics, Azure SQL DW, HDInsight/Databricks, NoSQL DB)
  • Developed JavaMap Reduce programsfor the analysis of sample log file stored in cluster.
  • Developed Simple to complex Map/reduce Jobs using HiveandPig
  • Developed Map Reduce Programs for data analysis and data cleaning.
  • Developed Spark scripts using Python on Azure HDInsight for Data Aggregation, Validation and verified its performance over MR jobs
  • Performed Data Visualization and Designed Dashboards with Tableau and generated complex reports including chars, summaries, and graphs to interpret the findings to the team and stakeholders.
  • Wrote documentation for each report including purpose, data source, column mapping, transformation, and user group.
  • Extracted and loaded data into Data Lake environment (MS Azure) by using Sqoop which was accessed by business users.
  • Primarily involved in Data Migration process using Azure by integrating with GitHub repository and Jenkins.
  • Extract Transform and Load data from Sources Systems to Azure Data Storage services using a combination of Azure Data Factory, T-SQL, Spark SQL and U-SQL Azure Data Lake Analytics. Data Ingestion to one or more Azure Services - (Azure Data Lake, Azure Storage, Azure SQL, Azure DW) and processing the data in inAzure Databricks.
  • Utilized the clinical data to generate features to describe the different illnesses by using LDA Topic Modelling.
  • Utilized Waterfall methodology for team and project management.
  • Used Git for version control with Data Engineer team and Data Scientists colleagues.

ENVIRONMENT: Ubuntu Hadoop, Spark (PySpark, Nifi, Jenkins, Kafka, Talend, SparkSQL, SparkMLIib), MS Azure, Azure Data Bricks, Azure SQL. Azure Data Factory (ADF), Azure Data Lake, Pig, Python, 3.x(NltkPandas), Tableau, GitHub, and OpenCV.

Confidential, Columbus, OH

Big Data Engineer

Responsibilities:

  • Maintain Hadoop, Hadoop ecosystems, and database(s) with updates/upgrades, performance tuning and monitoring
  • Worked in Agile environment this uses Jira to maintain the story points and Kanban model.
  • Created and Implemented highly scalable and reliable highly scalable and reliable distributed data design using NoSQL/Cassandra technology.
  • Developed multiple MapReduce jobs in java for data cleaning and preprocessing.
  • Extracted files from MySQL through Sqoop and placed in HDFS and processed.
  • Created and maintained Technical documentation for launching HADOOP Clusters and for executing Hive queries and Pig Scripts.
  • Worked on Implementation of a log producer in Scala that watches for application logs transform incremental log and sends them to a Kafka and Zookeeper based log collection platform
  • Maintaining different cluster security settings and involving in creation and termination of multiple cluster environment.
  • Worked with the Data Science team to gather requirements for various data mining projects.
  • Involved in creating Hive tables, and loading and analyzing data using hive queries.
  • Implemented Kafka High level consumers to get data frsom Kafka partitions and move into HDFS
  • Used Oozie and Zookeeper operational services for coordinating cluster and scheduling workflows.
  • Implementing MR programs to analyze large datasets in warehouse for business intelligence purpose
  • Developed customized Hive UDFs and UDAFs in Java, JDBC connectivity with hive development and execution of Pig scripts and Pig UDF's.
  • Developed Simple to complex MapReduce Jobs using Hive and Pig.
  • Worked on distributed frameworks such as Apache Spark and Presto in Amazon EMR, Redshift and interact with data in other AWS data stores such as Amazon 53 and Amazon DynamoDB.
  • Responsible for data gathering from multiple sources like Teradata, Oracle, Sql server etc.
  • Involved in running Hadoop jobs for processing millions of records of text data.
  • Worked with application teams to install operating system, Hadoop updates, patches, version upgrades as required.
  • Handling structured and unstructured data and applying ETL processes.
  • Wrote MapReduce jobs using Java API and Pig Latin.
  • Loaded the data from Teradata to HDFS using Teradata Hadoop connectors.
  • Used Amazon EMR to simplify big data processing and to manage Hadoop framework.
  • Wrote Pig scripts to run ETL jobs on the data in HDFS.
  • Used Hive to do analysis on the data and identify different correlations.
  • Worked on importing and exporting data from Oracle and DB2 into HDFS and HIVE using Sqoop.

ENVIRONMENT: Hadoop, HDFS, MapReduce, Unix, REST, Redshift, Python, Pig, Hive, HBase, Storm, NoSql, Flume, Zookeeper, Kibana, Cloudera, Hortonworks, SQL, Amazon Web Services, SAS, Vertica, Kafka, Cassandra, Informatica, Teradata, Scala, Spark Streaming, Spark.

Confidential

Data Engineer

Responsibilities:

  • Implemented Flume, Spark framework for real time data processing.
  • Developed simple to complex Map Reduce jobs using Hive and Pig for analyzing the data.
  • Optimized Map/Reduce Jobs to use HDFS efficiently by using various compression mechanisms.
  • Developed big data ingestion framework to process multi TB data including data quality checks, transformation, and stored as efficient storage formats like parquet and loaded into Amazon S3 using Spark Scala API and Spark.
  • Wrote the Spark code in Scala to connect to Hbase and read/write data to the HBase table.
  • Extracted data from different databases and to copy into HDFS using Sqoop and has an expertise in using compression techniques to optimize the data storage.
  • Implemented Kafka producers create custom partitions, configured brokers and implemented High level consumers to implement data platform.
  • Responsible for building scalable distributed data solutions using Hadoop and migrate legacy applications to Hadoop.
  • Scheduled automated tasks with Oozie for loading data into HDFS through Sqoop and pre-processing the data with Pig and Hive.
  • Worked on scalable distributed computing systems, software architecture, data structures and algorithms using Hadoop, Apache Spark and Apache Storm etc.
  • Ingested streaming data into Hadoop using Spark, Storm Framework and Scala.
  • Developed the technical strategy of using Apache Spark on Apache Mesos as a next generation, Big Data and "Fast Data" (Streaming) platform.
  • Created the Spark Streaming code to take the source files as input.
  • Used Oozie workflow to automate all the jobs.
  • Exported the analyzed data into relational databases using Sqoop for visualization and to generate reports for the BI team.
  • Developed Pig UDF's to know the customer behavior and Pig Latin scripts for processing the data in Hadoop.
  • Copied the data from HDFS to MongoDB using pig/Hive/Map reduce scripts and visualized the streaming processed data in Tableau dashboard.
  • Continuously monitored and managed the Hadoop Cluster using Cloudera Manager.

ENVIRONMENT: Spark, Kafka, Hadoop, AWS, Sqoop, HDFS, Oracle, SQL Server, MongoDB, Python, Scala, Shell Scripting, Tableau, Map Reduce, Oozie, Pig, Hive.

Confidential

Hadoop Developer

Responsibilities:

  • Experience in writing customMapReduceprograms &UDF's in Java to extendHiveandPigcore functionality.
  • Involved in collecting, aggregating and moving data from servers toHDFSusingFlume.
  • Involved in creatingOozieworkflow and Coordinator jobs to kick off the jobs on time for data availability.
  • Enabled speedy reviews and first mover advantages by usingOozieto automate data loading into the Hadoop Distributed File System andPIGto pre-process the data.
  • Developed job workflow inOozieto automate the tasks of loading the data intoHDFSand few otherHivejobs.
  • DevelopedPigScripts to store unstructured data inHDFS.
  • Responsible for coding Java Batch, Restful Service,MapReduceprogram, Hive query's, testing, debugging, Peer code review, troubleshooting and maintain status report.
  • Installed and configuredFlume,Hive,Pig,SqoopandOozieon the Hadoop cluster.
  • DevelopedFlumeAgents for loading and filtering the streaming data intoHDFS.
  • Handling continuous streaming data comes from different sources usingFlumeand set destination asHDFS.
  • UsedHiveto analyze the partitioned and bucketed data and compute various metrics for reporting.
  • Worked on various performance optimizations like using distributed cache for small datasets, partition and bucketing inHive, doing map side joins etc.
  • DevelopedPigLatin scripts to extract and filter relevant data from the web server output files to load into HDFS.
  • Analyzed the data by performingHivequeries and runningPigscripts to study customer behavior.
  • OptimizedMapReduceJobs to useHDFSefficiently by using various compression mechanisms.

ENVIRONMENT: Hadoop, HDFS, Oozie, HBase, Sqoop, RDBMS/DB, Flat files, MySQL, Java. Map Reduce, Hive, Pig, Flume.

We'd love your feedback!