Big Data Engineer Resume
Albany, NY
PROFESSIONAL SUMMARY:
- Experience in working with Hadoop Ecosystem, Big Data and Java/J2EE Technologies, Database Management Systems (DBMS)and Enterprise - level Cloud Base Computing and Applications.
- Excellent understanding of Distributedstorage systems like Hadoop Distributed File System (HDFS) and Batchprocessing systems like Map Reduce and Yarn.
- Comprehensive experience on Hadoop eco system and its technologies such as MapReduce , Hive , Pig , HBase , Spark , Sqoop , Flume , Kafka , Oozie , Zookeeper and Impala with CDH distributions and EC2 cloud computing with AWS.
- Expertise in depth understanding/knowledge of Hadoop Architecture and various components such as HDFS, Name Node, MapReduce, Data Node, Resource Manager, Node Manager,Job Tracker, Task Tracker,MRv1 and MRv2 (YARN).
- Experienced in optimization of MapReduce algorithm using combiners and practitioners for analyzing the big data as per the requirement to deliver the best results.
- Experienced with data architecture including data ingestion pipeline design, Hadoop information architecture, data modeling and data mining and advanceddata processing.
- Created Partitions and Bucketing concepts in Hive and designed both Managed and External tables in Hive to optimize performance.
- Good knowledge in using ofPig as ETL tool to do transformations, eventjoins, filter and some pre-aggregation.
- Working experience with DevelopingUser Defined Functions (UDFs), UDTF and UDAF for Apache Pig and Hive using Python and Java languages to analyze data.
- Experienced in performing in memory dataprocessing for batch, real-time , and advancedanalytics using Apache Spark (Spark Core, Spark SQL&Spark-Shell) , Streaming.
- Experienced in migrating Map Reduce programs into SparkRDD transformations and actions to improveperformance.
- Experienced with SparkContext, Spark-SQL, DataFrame, PairRDD's, SparkYARNand build tools like Maven and SBT.
- Good knowledge of integrating SparkStreaming with Kafka for real time processing of streaming data
- Experience in designing and implementationof data ingestion patterns using Kafka and Spark, Developing Custom analytics using Hadoop, Spark , Spark Streaming , Spark SQL , Spark and Scala.
- Experience in loadinglogs from multiple sources directly into HDFS using Kafka.
- Scheduling, monitoring job workflows , identifying failures with Oozie , and integrating jobs with Zookeeper .
- Knowledge on Integrating and managing APIs exposing MicroServices ( REST , SOAP ) including development and support of Java services.
- Experience in implementing Microservices, RESTful APIs, and event driven architectures
- Ingested data into Hadoop from various data sources like Oracle, MySQL and Teradata using Sqoop tool. Created Sqoop job with incremental load to populate Hive External tables.
- Exported the analyzed data to the relational databases using SQOOP for visualization and to generate reports for the BI team.
- Expertise in AmazonAWS concepts like Lambda, S3, EBS, RDS, Dynamo DB, Red shift, SQS, SNS, EMR&EC2 web services which providefast and efficient processing of Big Data.
- Strong knowledge in NOSQL column oriented databases like HBase, Cassandra and their integration with Hadoop cluster.
- Good Experience onscheduling and monitoring tools like Oozie&Zookeeper
- Hands on Evaluation of ETL (Talend) and OLAP tools and recommend the most suitable solutions based on business requirements.
- Used advanced analytics and search engine ELK stack (Elasticsearch, LogstashandKibana) to process large datasets and visualize the results based on aggregations and filters on structured and unstructured fields
- Experience in running Hadoop streaming jobs to process terabytes of data working with file formats like Parquet, Avro, RCFile, SequenceFiles,ORC and JSONRecord, DAT, RC, ORC,CSV etc.
- Good knowledge and working experience in using SBT and Maven Scripts for building and deploying the application in web/App servers and worked on secured FTP, SCP Process.
- Worked on various Hadoop Distributions (Cloudera (CDH4, CDH5), MapR, Hortonworks Distributions (HDP)and AmazonAWSto implement and make use of those.
- Experience in development of applications using Java/ JDK, J2EE(TM) Technology- J2EE technologies (JDBC) and web Technologies like HTML, CSS.
- Experienced in backend development using SQL, stored procedures on Rational Database Management System (RDBMS).
- Worked on various Tools and IDEs like Eclipse,Netbeans, Visio, Apache Ant-Build Tool, MS-Office, PLSQL Developer, SQL*Plus
TECHNICAL SKILLS:
Operating Systems: Linux, MacOS, Windows 10,Windows 8, Windows
Hadoop/Big Data: HDFS, MapReduce, Pig, Hive, ImpalaSparkSQL, HBase, Kafka, Sqoop, Spark Streaming, PySpark, Oozie, Zookeeper,Hue, Ambari, Strom, Scala.
Hadoop Distribution: HortonWorks Distributions (HDP), Cloudera (CDH4,CDH5), MapR/Map Reduce, Amazon AWS.
Programming/Scripting languages: Java, Linux shell scripts, Python, Scala, JavaScript, HTML, XML.
Database: MySql,PL/SQL,SQL Developer,Cassandra, Couch DB, Teradata, HBase
ETL: Talend
Real Time/Stream processing: Apache Storm,Apache Spark
Build tools: Ant, Maven, SBT
Cloud: AWS, Microsoft Azure, Google Cloud
Java Technologies: JSP, Servlets, Struts, Spring and Hibernate, JavaBeans, JDBC.
IDE's, Web/App servers: Intellej, Eclipse, NetBeans, Web Logic, Web Sphere, Tomcat.
WORK EXPERIENCE:
Confidential, Albany, NY
Big Data Engineer
Responsibilities:
- Worked on analyzing Hadoop cluster using different big data analytic tools including Spark (Spark SQL, Spark-Shell), Kafka,Hive and MapReduce.
- Implemented Spark using Scala and utilizing Data frames and Spark SQL API for faster processing of data.
- Developed Spark code using Scala and SparkSQL /Streaming for faster testing and processing of data and developed very quick poc's on Spark in the initial stages of the product.
- Involved in data ingestion into HDFS using Sqoopfrom variety of sources using the connectors like JDBC and import parameters.
- Involved in Developing Hive scripts to parse the raw data, populate staging tables and store the refined data in partitioned tables in the Hive.
- Developed Hive Scripts (HQL) for automating the joins for different sources.
- Developed Hive UDFs (User Defined Functions) where the functionality is too complex, using Python and Javalanguages.
- Automated all the jobs from pulling data from databases to loading data into SQL server using shell scripts.
- Created and worked Sqoop jobs with incremental load to populate Hive External tables.
- Performed various data warehousing operations like de-normalization and aggregation on Hive using DML statements.
- Written Map/Reduce programs, Hive UDFs to specify the conditions to separate the fraudulent claims .
- Experienced in Hive to do transformations, event joins, filter bot traffic and some pre-aggregations before storing the data into HDFS.
- Used maven and SBT to build and deploy the Jars for MapReduce, Pig and Hive UDFs.
- Developing Hive UDFs for change data capture and deltarecord processing between newly arrived data and already existing data in HDFS.
- Imported and extracted the needed data using Scoop from the server into HDFS and BulkLoaded the cleaned data into HBase using MapReduce .
- Involved in loading data into HBaseusing HBase Shell, HBaseClientAPI, Pig and Sqoop.
- Worked with relational database systems (RDBMS) such as MySQL, MSSQLand NoSQL database systems like HBase and Cassandra.
- Implemented Data Pipe Lines using Kafka - Stream Data Platform that captures Stream events in Topics and feeds data to Data systems such as HDFS , also cleanses and aggregates data on the fly to channelize the data to Data Lake .
- Pre-processed large sets of structured and semi-structured data, with different formats like Text Files, Avro, Sequence Files, and JSON Record.
- Implemented Oozieworkflows for MapReduce, Hive and sqoop actions.
- Worked with Oozie and Zookeeper to manage the flow of jobs and coordination in the cluster.
Environment: Spark,HDFS, Pig, Hive, MapReduce, Scala, Sqoop, SparkSQL, Kafka, Spark, Python, Linux Shell Scripting, JDBC,Git.
Confidential, Minneapolis, MN
Big Data Engineer
Responsibilities:
- Developed Pig Scripts, PigUDFs,HiveScripts, HiveUDFs, Python Scripting and used Spark(SparkSQL, Spark-shell)to process data inHortonworks.
- Performed advanced procedures like text analytics and processing using the in-memory computing capabilities of Spark.
- Designed and Developed Scala workflows for data pull from cloud based systems and applying transformations on it.
- Usage of Sqoopto importdata into HDFS from MySQL database and vice-versa.
- Implemented optimized joins to perform analysis on different data sets using MapReduce programs.
- Experienced in optimizing shuffle and sort phase in MapReduce Phase.
- Experience in processing of load and transform the large data sets of structured , unstructured and semistructured data in Hortonworks.
- Implemented Partitioning, Dynamic Partitions and Buckets in HIVE & Impalafor efficient data access.
- Extensively worked on HiveQL, join operations, writing customUDF's and having good experience in optimizing Hive Queries.
- Experienced in running query using Impala and used BI tools and reporting tool ( tableau ) to run ad-hoc queries directly on Hadoop.
- Created Impalaviews over HIVE tables and used them to extract the data for Portal Processing and Implemented Impala for data analysis.
- Tested Apache Tez, an extensible framework for building high performance batch and interactive data processing applications, on Pig and Hive jobs
- Experience in using Spark framework with Scalaand Python. Good exposure to performance tuning hive queries and MapReduce jobs in spark(SparkSQL) framework on Hortonworks.
- DevelopedScala & Python scripts, UDF's using both Data frames/SQL and RDD/MapReduce in SparkSQL for DataAggregation, queries and writing data back into RDBMS through Sqoop.
- Configured Spark streaming to receive real time data from the Kafka and store the streamdata to HDFS using Scala and Python.
- Collect the data using Spark streaming and dump into HBase and Cassandra.Used the Spark- Cassandra Connector to load data to and from Cassandra.
- Configured & deployed and maintained multi-node Dev and Test Kafka Clusters . Developed multiple Kafka Producers and Consumers from scratch as per the business requirements.
- Developing Java Restful web-services to make it work with HBase .
- Used Hive to analyze data ingested into HBase by using Hive-HBase integration and compute various metrics for reporting on the dashboard.
- Involved in HBASE setup and storing data into HBASE, which will be used for analysis and Fetch data to/from HBase using MapReduce jobs on Hortonworks.
- Created ETL Scripts for Dataacquisition and Transformation using Talend, Extensively worked on the ETL mappings, analysis and documentation of OLAP reports .
- Working on ETL data flow development using Hive/Pig scripts and loading from Hive views .
- Collecting and aggregating large amounts of log data using Kafka and staging data in HDFS for further analysis
- Provision, monitor and maintain AWS EC2 instances , watching the security and manage the AWS S3 bucket storage on AWS cloud environment.
- Experience in large scale data migration of LINUX Logical Volumes (LVM) and restore the data using AWS snapshots .
- Developed a Spark job in Java which indexes data into Elasticsearch from external Hive tables which are in HDFS.
- Knowledge on application deployment as a Micro Services and integrate them with Kafka and Elastic search.
- Used ELK ( Elasticsearch , Logstash and Kibana ) stack for log analysis, monitoring, alerts and visualization.
- Created and defined job work flows as per their dependencies in Oozie and e- mail notification service upon completion of job for the particular team that request for the data and monitoredjobs using Oozie on Hortonworks.
- Experience in designing both time driven and datadriven automated workflows using Oozie.
Environment: Hadoop (MAPR), HDFS, Python Scripting, Cassandra, Map Reduce, Hive,Impala, Pig, SparkSQL, Spark Streaming, Spark, Sqoop, WebSphere, Oozie, REST Web Services, DB2, AWS S3, Java, JDBC, Python, Scala, Tableau, UNIX Shell Scripting,Git, Elastic Search, Kibana.
Confidential, Detroit, MI
Big Data Engineer
Responsibilities:
- Involved in the high-level design of the Hadoop architecture for the existingdata structure and Problem statement and setup the 64-node cluster and configured the entire Hadoop platform.
- Implemented Data Interface to get information of customers using RestAPI and Pre - Process data using MapReduce and store into HDFS( Hortonworks).
- Extracted files from MySQL, Oracle, and Teradata through Sqoop and placed in HDFS Cloudera Distribution and processed.
- Configured Hive metastore, which stores the metadata for Hive tables and partitions in a relational database.
- Worked with various HDFS file formats like Avro, Sequence File, Jsonand various compression formats like Snappy, bzip2.
- Developed efficient MapReduce programs for filtering out the unstructured data and developed multiple MapReduce jobs to perform datacleaning and preprocessing on Hortonworks.
- Developed the Pig UDF's to pre-process the data for analysis and Migrated ETL operations into Hadoop system using Pig Latin scripts and Python Scripts.
- Used Pig as ETL tool to do transformations, event joins, filtering and some pre-aggregations before storing the data into HDFS.
- Troubleshooting, debugging & altering Talend issues , while maintaining the health and performance of the ETL environment .
- Developed Hive queries for data sampling and analysis to the analysts.
- Loaded data into the cluster from dynamically generated files usingFlume and from relationaldatabase management systems using Sqoop.
- Developed custom Unix SHELL scripts to do pre and post validations of master and slave nodes, before and after configuring the name node and datanodes respectively.
- Experienced in running Hadoop streaming jobs to process terabytes of formatted data using Python scripts.
- Developed small distributed applications in our projects using Zookeeper and scheduled the workflows using Oozie.
- Developed complex Talend jobs mappings to load the data from various sources using different components.
- Developed a SCP Stimulator which emulates the behavior of intelligent networking and Interacts with SSF.
- Created HBase tables from Hive and Wrote HiveQL statements to access HBase table’s data.
- Proficient in designing Row keys and Schema Design for NoSQL Database Hbase and knowledge of other NOSQL database Cassandra .
- Used Hive to perform data validation on the data ingested using scoop and flume and the cleansed data set is pushed into Hbase .
- Created a MapReduce program which looks into data in HBasecurrent and prior versions to identify transactional updates. These updates are loaded into Hive externaltables which are in turn referred by Hive scripts in transactionalfeeds generation.
Environment: Hadoop (Cloudera), HDFS, Map Reduce, Hive, Scala,Python,Pig, Sqoop, WebSphere, Hibernate, spring, Oozie, REST Web Services, AWS, Solaris, DB2, UNIX Shell Scripting, JDBC.
Confidential
Java Developer
Responsibilities:
- Involved in complete software development lifecycle with object oriented approach of clients business process and continuous client feedback.
- Worked on designing and developing a complete service oriented systembased on SOA principles and architecture in agile development environment.
- Actively participated in Object Oriented Analysis Design sessions of the Project, which is based on MVC Architecture using Spring Framework.
- Used Core JAVA Collection API, Generics, Annotations, Reflection API, multi-threading in application development.
- Developed user interfaces using JSP and form beans with JavaScript to reduce round trips to the server.
- Used various Core Java concepts such as Exception Handling , Collection APIs to implement various features and enhancements.
- Involved in Java, J2ee, Spring 4.0, Restful WebServicesand WebSphere5.0/6.0 in a fast paced development environment.
- Involved in development of controller component using Servlets and view component using JSP, XSLT, CSS, HTML and JavaScript for the client side validation.
- Designed and implemented UI layer using JSP, JavaScript, XML, XHTML, XSL, XSLT and business logic using Servlets, JSP and J2EE framework.
- Developed the UI components using JQuery and JavaScript Functionalities.
- Create, edit and maintain sites implementing responsive design & themes using front end development frameworks including Bootstrap .
- Utilized AngularJS along with Sass/ Bootstrap to develop responsive web pages and NotifyJS, a JQuery plugin for user notifications
- Involved on the back end to modifybusiness logic by making enhancements.
- Doing coding based on MVC architecture using Core Java - java .util.*, Collections , Hashmaps , Multithreading , strings , LinearDatastructures like arrays & Lists and HeapTrees etc, J2EE tools and Swings along with all OOP's concepts and security features.
- Involved in developing the UI pages using HTML , DHTML , CSS , JavaScript , JSON , JQuery and Ajax .
- Implemented design patterns like singleton , DAO of COREJAVA in developing applications.
- Involved in the design and implementation of the architecture for the project using OOAD , UML design patterns.
- Involved in design and implementation of contractWeb service.
- Involved in the business logic-coding framework to seamlessly map the business logic into respective value beans.
- Involved in publishing the web services to help users interacting with web services.
- Worked closely with requirements to translate businessrules into businesscomponent modules.
- Involved in the migration of independent parts of the system to use persistence technology such as JDBC.
- Used Eclipse as IDE for development, build, deployment and testing the application.
- Created PL/SQL packages , Stored Procedures, Functions, Triggers and Views to retrieve, manipulate and migrate complex data sets in MySQL Databases.
- Wrote database queries using SQL and PL/SQL for accessing, manipulating and updating MySQL database.
- Accessed data from database by writing SQL statements in Java programs using JDBC connectivity.
- Used Clear Case to mergecode and deploythem in to a centraldepository location.
- Multithreading used to enhance interaction between ratematrix and ECM systems.
Environment: JDK 1.7, Junit, JDBC,MySQL, JSP, UML, JUNIT, Angular.js, HTML, CSS,Bootstrap, JQuery Hibernate 4.0, Spring, Struts, WebSphere 5.0/6.0, SOAP, Web Services (SOAP, RESTFUL 4.0), Java Script, LDAP.
