Sr. Bigdata/hadoop Developer Resume
Fort Worth, TX
SUMMARY
- Over 8 years of professional IT experience including 3+ years in Big Data ecosystem related technologies. Expertise in Big Data technologies a consultant, proven capability in project based team work and also as an individual developer with good communications skills.
- Hands - on experience with major components in Hadoop Ecosystem like Map Reduce, HDFS, YARN, Hive, Pig, HBase, Sqoop, Oozie, Cassandra, Impala and Flume.
- Knowledge in installing, configuring, and using Hadoop ecosystem components like Hadoop Map Reduce, HDFS, HBase, Oozie, Hive, Sqoop, Pig, spark, kafka, storm, Zookeeper and Flume.
- Experience in installation, configuration, supporting and monitoring Hadoop clusters using Apache, Cloudera distributions and AWS.
- Experience with new Hadoop 2.0 architecture YARN and developing YARN Applications on it.
- Experience with Apache Spark’s Core, Spark SQL, Spark Streaming.
- Experience with distributed systems, large-scale non-relational data stores and multi-terabyte data warehouses.
- Hands on experience working on Java to implement MAPREDUCE jobs.
- Experience on Performance Tuning to ensure that assigned systems were patched, configured and optimized for maximum functionality and availability. Implemented solutions that reduced single points of failure and improved system uptime to 99.9% availability.
- Firm grip on data modeling, database performance tuning and NoSQL map-reduce systems.
- Solid experience in writing complex SQL queries. Also, experienced in working with NOSQL databases like Cassandra.
- Excellent knowledge on Hadoop Architecture and ecosystems such as HDFS, Hive, Pig, Sqoop, Job Tracker, Task Tracker, Name Node, Data.
- Experience on Data Virtualization tools like Tableau.
- Hands on experience working on version control GIT.
- Hands on experience in Agile and scrum methodologies.
- Performed importing and exporting data into HDFS and hive using Scoop.
- Experience in managing HBase database and using it to update/modify the data.
- Experience in running MapReduce and Spark jobs over YARN.
- Experience with Cloudera distributions (i.e.) CDH3/CDH4.
- Extending Hive and PIG core functionality by writing UDFs.
- Experience in Object Oriented Analysis and Design (OOAD) and development of software using UML Methodology.
- Excellent experience in developing User Interface/Front End applications using technologies such as HTML, HTML5, CSS, CSS3, JavaScript and Angular.JS.
- Experience with SSIS and T-SQL stored procedures to transfer data from OLTP databases to staging area and finally transfer into data marts.
- Hands-on experience in complete project life cycle (design, development, testing and implementation) of Client Server and Web applications.
- Excellent interpersonal skills, good experience in interacting with clients with good team player and problem solving skills.
- Strong team player, ability to work independently and in a team as well, ability to adapt to a rapidly changing environment, commitment towards learning.
TECHNICAL SKILLS
Big Data: HDFS, MapReduce, Hive, Pig, Zookeeper, Apache Spark, Core, Yarn, Spark SQL and Data frames, Scala, Ambari
Utilities: Sqoop, Flume, Kafka, Oozie
No SQL Databases: Hbase,Cassandra
Languages: C, C++, Java,Python,J2EE, PL/SQL, MR, Pig Latin, HiveQL, Unix shell scripting and Scala
Operating Systems: Sun Solaris, RedHat Linux, Ubuntu Linux and Windows XP/Vista/7/8, Suse Linux
Web Technologies: HTML, DHTML, XML, HTML5, CSS.
Databases and Data Warehousing: Teradata, DB2, Oracle 9i/10g/11g, SQL Server, MySQL
Tools: and IDE: Maven, Toad, Eclipse, NetBeans, Sonar, JDeveloper, DB Visualizer, Tableau, Talend
Methodologies: Agile Software Development, waterfall
PROFESSIONAL EXPERIENCE
Confidential, Fort Worth, TX
Sr. BigData/Hadoop Developer
Responsibilities:
- Worked on a live90 nodes Hadoop clusterrunningCDH4.4.
- Extracted the data from Teradata into HDFS usingSqoop.
- Workedwith Sqoop (version 1.4.3)jobs with incremental loadto populate Hive External tables.
- Involved in writingPig (version 0.10)scripts to transform raw datafrom several data sources into forming baseline data.
- DevelopedHive(version 0.10) scripts for end user / analyst requirements to perform ad hoc analysis.
- Developed ETL data pipelines using Spark, Spark streaming and Scala.
- Loaded data from RDBMS to Hadoop using Sqoop.
- Worked with on Java for implementing MapReduce Jobs.
- Worked collaboratively to manage build outs of large data clusters and real time streaming with Spark.
- Responsible for loading Data pipelines from web servers and Teradata using Sqoop with Kafka and Spark Streaming API.
- Actively involved in architecture of DevOps platform and cloud solutions.
- Involved in supporting data analysis projects using Elastic Map Reduce on the Amazon Web Services (AWS) cloud.
- Developed multiple map Reduce jobs in Java and Python for data cleaning and processing.
- Used Spark for interactive queries, processing of streaming data and integration with popular NoSQL database for huge volume of data.
- Performance tuning the Spark jobs by changing the configuration properties and using broadcast variables.
- Worked on performing transformations & actions on RDDs and Spark streaming data.
- Responsible in handling Streaming data from web server console logs.
- Solved performance issuesin Hive and Pig scripts with understanding of Joins, Group and aggregation and how does it translate to MapReduce jobs.
- DevelopedUDFsin Java as and when necessary to use in PIG and HIVE queries.
- Involved in collecting metrics for Hadoop clusters usingGanglia and Ambari.
- Installed Ambari on existing Hadoop cluster.
- Solved performance issues in Hive and Pig scripts with understanding of Joins, Group and aggregation and how does it translate to MapReduce jobs Experience in usingSequence files, RCFile, AVRO and HARfile formats.
- DevelopedOozieworkflow for scheduling and orchestrating the ETL process
- Worked with the admin team in designing and upgrading CDH 3 to CDH 4 Environment Cluster
- Worked with both MapReduce 1 (Job Tracker) and MapReduce 2 (YARN) setups
- Involved in monitoring and managing the Hadoop cluster usingCloudera Manager.
- Implemented best income logic using Pig scripts and UDFS.
- Extracted files from CouchDB through Sqoop and placed in HDFS and processed.
- Worked on Hive for exposing data for further analysis and for generating transforming files from different analytical formats to text files.
- Worked on version control GIT.
- Imported data fromMySQLserver and other relational databases to Apache Hadoop with the help ofApache Sqoop.
- CreatingHive tablesand working on them for data analysis in order to meet the business requirements.
- Implemented a script to transmit sys print information from Oracle to HBase using Sqoop.
- Responsible for building scalable distributed data solutions using Hadoop.
- Evaluated business requirements and prepared detailed specification’s that follow project guidelines required to develop written programs.
- Worked with professional software engineering practices for the full software development life cycle including coding standards, source control management control and build processes.
Environment: Hadoop, BigData, MapReduce, HDFS, Hive, HBase, Spark, Python, AWS, DevOps, Scala, Sqoop, Pig, Flume, Oracle 11/10g, DB2, Teradata, MySQL, Eclipse, PL/SQL, Java, Linux, Shell Scripting, SQL Developer, SOLR, Ambari.
Confidential, Woodland Hills, CA
Bigdata/Hadoop Developer
Responsibilities:
- Worked on a live90 nodes Hadoop clusterrunningCDH4.1.
- Developed hive queries on data logs to perform a trend analysis of user behavior on various online modules.
- Developed the Pig UDF'S to pre-process the data for analysis.
- Developed use case diagrams, class diagrams, database tables, and mapping between relational database tables and object oriented Java objects using Hibernate.
- Involved in the setup and deployment of Hadoop cluster.
- Developed Map Reduce programs for some refined queries on big data.
- Involved in loading data from UNIX file system to HDFS.
- Loaded data into HDFS and extracted the data from MySQL into HDFS using Sqoop.
- Exported the analyzed data to the relational databases using Sqoop and generated reports for the BI team.
- Managing and scheduling jobs on a Hadoop cluster using Oozie.
- Along with the Infrastructure team, involved in design and developed Kafka and Storm based data pipeline.
- Designed and configured Kafka cluster to accommodate heavy throughput of 1 million messages per second. Used Kafka producer 0.6.3 API's to produce messages.
- Installed, Configured Talend ETL on single and multi-server environments.
- Troubleshooting, debugging & fixing Talendspecific issues, while maintaining the health and performance of the ETL environment.
- Worked for DevOps Platform team responsible for specialization areas related to Chef for Cloud Automation.
- Worked on Amazon Web Services (AWS) to complete set of infrastructure and application services.
- Developed Merge jobs in Pythonto extract and load data into MySQL database.
- Involved in unit testing in SSIS Packages, Query verification in SSIS & stored procedures.
- Created and modified several UNIX shellScripts according to the changing needs of the project and client requirements. Developed UNIX shellscripts to call Oracle PL/SQL packages and contributed to standard framework.
- Developed Simple to complex Map/reduce Jobs using Hive.
- Implemented Partitioning and bucketing in Hive.
- Mentored analyst and test team for writing Hive Queries.
- Involved in setting up of HBase to use HDFS.
- Extensively used Pig for data cleansing.
- Setup SOLR and configured multi-Core.
- Loaded streaming log data from various Webserver into HDFS using Flume.
- Along with the Infrastructure team, involved in design and developed Kafka and Storm based Performed benchmarking of the No-SQL databases, Cassandra and Hbase streams.
- Implemented Spark using Scala and SparkSQL for faster testing and processing of data.
- Worked with Spark and Scala mainly in framework exploration for transition from Hadoop/MapReduce to Spark.
- Supported in setting up QA environment and updating configurations for implementing scripts With Pig and Sqoop.
- Involved in collecting and aggregating large amounts of log data using Apache Flumeand staging data in HDFS for further analysis.
- Configured Flume to extract the data from the web server output files to load into HDFS.
- Developed workflow in Oozie to automate the tasks of loading the data into HDFS and pre-processing with Pig.
Environment: Unix Shell Scripting, Python, Oracle 11g, HDFS, Kafka, Storm, AWS, Spark, DevOps, ETL, Java, Pig, Linux, Cassandra, MapReduce, Ms Access, Toad, SQL, Scala, MySQL Workbench, XML, No-SQL, MapReduce, SOLR, HBase, Hive, Sqoop, Flume, Talend, Oozie.
Confidential, Boston, MA
Hadoop Developer
Responsibilities:
- Worked on a live 110 nodes Hadoop cluster running CDH4.0.
- Processed data into HDFS by developing solutions analyzed the data using MapReduce, pig Hive and produce summary results from Hadoop to downstream systems.
- Involved in development of MapReduce job using HiveQL Statements.
- Work closely with various levels of individuals to coordinate and prioritize multiple projects, estimate scope, schedule and track projects throughout SDLC.
- Worked in Hadoop MapReduce and HDFS. Developed multiple MapReduce jobs in java for data cleaning and processing.
- Involved in MapReduce job using HiveQL query for data stored in HDFS.
- Written Storm topology to accept the events from Kafka producer and emit into Cassandra DB.
- Experienced in managing and reviewing Hadoop Log files.
- Designed a data warehouse using Hive.
- Handling structured, semi structured and unstructured data.
- Developed simple to complex MapReduce jobs using Hive and Pig.
- Extensively used pig for data cleansing.
- Created partitioned tables in Hive.
- Managed and reviewed Hadoop log files.
- Cluster coordinating services through ZooKeeper.
- Mentored analyst and test team for writing Hive Queries.
- Designed and configured Kafka cluster to accommodate heavy throughput of 1 million messages per second. Used Kafka producer 0.6.3 API's to produce messages.
- Troubleshooting, debugging & fixing Talendspecific issues, while maintaining the health and performance of the ETL environment.
- Developed Simple to complex Map/reduce Jobs using Hive.
- Implemented Partitioning and bucketing in Hive.
- Mentored analyst and test team for writing Hive Queries.
- Involved in setting up of HBase to use HDFS.
- Supports and assists QA Engineers in understanding, Testing and troubleshooting.
- Exported and analyzed data to the relational databases using Sqoop for visualization and to generate reports for the BI team.
- Developed workflow in Oozie to automate the tasks of loading the data into HDFS and Pre-processing with Pig.
- Involved in the database migrations to transfers to migrations to transfer data from one database to other and other and complete virtualization of many client application.
- Developed the pig UDF’S to pre-process the data for analysis.
- Created HBase tables to store variable data formats of data coming from different portfolios.
- Used Sqoop widely in order to import data from various systems/sources (like MYSQL) into HDFS.
Environment: Hadoop, MapReduce, HDFS, Hive, HBase, Sqoop, Pig, Teradata, MySQL, Shell Scripting, Kafka, Cassandra.
Confidential, Knoxville, TN
Java Developer
Responsibilities:
- Involved in gathering business requirements, analyzing the project and created UML diagrams such as Use Cases, Class Diagrams, Sequence Diagrams and flowcharts for the optimization Module using Microsoft Visio.
- Designed and developed Optimization UI screens for Rate Structure, Operating Cost, Temperature and Predicted loads using JSF myfaces, JSP, JavaScript and HTML.
- Configured faces-config.xml for the page navigation rules and created managed and backing beans for the Optimization module.
- Developed JSP web pages for rate Structure and Operating cost using JSF HTML and JSF CORE tags library.
- Designed and developed the framework for the IMAT application implementing all the six phases of JSF life cycle and wrote Ant build, deployment scripts to package and deploy on JBoss application server.
- Designed and developed Simulated annealing algorithm to generate random Optimization schedules and developed neural networks for the CHP system using Session Beans.
- Integrated EJB 3.0 with JSF and managed application state management, business process management (BPM) using JBoss Seam.
- Wrote Angular.JS controllers, views, and services for new website features.
- Developed Cost function to calculate the total cost for each CHP Optimization schedule generated by the Simulated Annealing algorithm using EJBs.
- Implemented spring web flow for the Diagnostics Module to define page flows with actions and views and created POJOs and used annotations to map them to SQL Server database using EJB.
- Wrote DAO classes, EJB 3.0 QL queries for Optimization schedule and CHP data retrievals from SQL Server database.
- Used Eclipse as IDE tool to develop the application and JIRA for bug and issue tracking
- Created combined deployment descriptors using XML for all the session and entity beans.
- Wrote JSF and JavaScript validations to validate data on the UI for Optimization and Diagnostics and Developed Web Services to have access to the external system (WCC) for the
- Designed and coded application components in an agile environment utilizing a test driven development approach.
- Skilled in test driven development and agile development.
- Created technical design document for the Diagnostics Module and Optimization module covering Cost function and Simulated Annealing approach.
- Involved in code reviews and performed version guidelines.
Environment: Java 1.5, J2EE, Microsoft Vision, EJB 3.0, JSP, JSF, JBoss Seam, JIRA, Web Services, JMS, JavaScript, Angular.js, HTML, ANT, Agile, JUnit, JBoss 4.2.2, MS SQL Server 2005, My ECLIPSE 6.0.1.
Confidential, Clearwater, FL.
Java Developer
Responsibilities:
- Analysis of system requirements and development of design documents.
- Involved in various client implementations.
- Development of Spring Services.
- Development of persistence classes using Hibernate framework.
- Development of SOA services using Apache Axis web service framework.
- Development of user interface using Apache Struts2.0, JSPs, Servlets, JQuery and Java Script.
- Developed client functionality using ExtJS.
- Development of JUnit test cases to test business components.
- Extensively used Java Collection API to improve application quality and performance.
- Vastly used Java 5 features like Generics, enhanced for loop, type safe etc.
- Providing production support and enhancements design to the existing product.
Environment: s: Java, SOA, spring, ExtJS, Struts 2.0, Servlets, JSP, GWT, JQuery, JavaScript, CSS, Web Services, XML, Oracle, Angular.JS and Windows.
