Hadoop / Spark Developer Resume
Round Rock, TX
PROFESSIONAL SUMMARY:
- 8+ years of IT experience in Analysis, Design, Development, Implementation, Integration and testing of Application Software in web - based environments, distributed n-tier products and Client/Server architectures.
- 5+ years of Big Data/Hadoop experience using Hadoop stack (MapReduce Programming, Pig, Hive, Oozie and Sqoop, Flume, Spark/Spark SQL, Strom, Kafka, Yarn/MRv2) and NoSQL technologies such as Cassandra, Hbase and MongoDB.
- Good knowledge of Hadoop Architecture and various components such as HDFS, Job Tracker, Task Tracker, Name Node, Data Node, Resource manager and YARN concepts.
- Sophisticated experience with Hadoop distributions like Apache, IBM Big Insights, Cloudera, Hortonworks&MapR.
- Expert in importing data using Sqoop into HDFS from various Relational Database Systems.
- Worked on standards and proof of concept in support of CDH4 and CDH5 implementation using AWScloud infrastructure.
- Hands on experience in implementing complex business logic and optimizing the query using Hive QL and controlling the data distribution by partitioning and bucketing techniques to enhance performance.
- Hands on experience in working with ClouderaCDH3 and CDH4 platforms
- Good knowledge in Architecting, Designing, re-Engineering and Performance Optimization.
- Experience with Oozie Workflow Engine in running workflow jobs with actions that run Hadoop Map Reduce and Pig jobs.
- Used Talend Open Studio for Big Data Integration and ETL operations.
- Experienced in creating and analyzing Software Requirement Specifications (SRS) and Functional Specification Document (FSD).
- Worked extensively with Dimensional modeling, Data migration, Data cleansing, Data profiling, and ETL Processes features for data warehouses.
- Expertise in workflow management tools like SQL Workbench, SQL Developer and TOAD tool for accessing the Database server.
- Very Good Knowledge in Logicaland Physical Data modeling, creating new data models, data flows and data dictionaries.
- Experience in understanding the security requirements for Hadoop and integrate with Kerberos authentication and authorization infrastructure.
- Experience in importing and exporting the data using Sqoop from HDFS to Relational Database systems/mainframe and vice-versa.
- Hands-on experience in using relational databases like Oracle, MySQL and MS-SQL Server.
- Experienced the integration of various data sources like Java, RDBMS, Shell Scripting, Spreadsheets, and Text files.
- Experience in Web Services using XML, HTML and SOAP.
- Excellent Java development skills using Java 6/7/8, J2EE, J2SE, Servlets, Junit, JSP, JDBC.
- Experience using middleware architecture using Sun Java technologies like J2EE, JSP, Servlets, and application servers like WebSphere and Web logic.
- Expertise in creating tables and distributing data in Hive by implementing Partitioning and Bucketing, writing complex scripts and User Defined Functions in Pig & Hive and customized MapReduce jobs in Java.
- Experience in working with Spark tools like RDD transformations, Spark MLlib and spark SQL
- Experiencein working with NoSQL database HBase, MongoDB and Cassandra
- Specialized in implementing Spark and Scala applications using higher order functions for both batch and interactive analysis.
- Experience in working withAmazon AWS services such as EMR, EC2, S3, etc.
- Worked and learned a great deal from Amazon Web Services (AWS) Cloud services like EC2, S3, EBS, RDS and VPC.
- Ability to spin up different AWS instances including EC2-classic and EC2-VPC using cloud formation templates.
- Expertise in using IDE like Eclipse, WebSphere (WSAD), NetBeans, My Eclipse, WebLogic Workshop
- Experience in developing and designing Web Services (SOAP and Restful Web services)
- Experience in developing Web Interface using Servlets, JSP and Custom Tag Libraries
- Experience in developing applications using SCRUM methodology and Agile Methodology
- Familiarity working with popular frameworks likes Struts, Hibernate, Spring MVC and AJAX.
- Exceptional ability to learn and master new technologies and to deliver outputs in short deadlines.
- Quick Learner, with high degree of passion and commitment in work.
- Proficiency with mentoring and on-boarding new engineers who are not proficient in Hadoop and getting them up to speed quickly.
- A good team player with excellent time management skills having profound insight to determine priorities, schedule work and meet critical deadlines.
TECHNICAL SKILLS:
Hadoop/Big Data: Apache Hadoop, MapReduce, Pig, Hive, Sqoop, Oozie, Flume, Zookeeper, Impala, Spark, Scala, Ambari, Impala, Kafka, YARN, HDFS, Ranger, Mapreduce, HBase, SOLR, Flume, MongoDB, Puppet, Oozie, Zookeeper, Spark, Talend, Hortonworks & Cloudera distributions
NOSQL Databases: HBase, MongoDB, Cassandra
Development Framework/IDE: RAD 8.x/7.x/6.0, IBM WebSphere Integration Developer 6.1, WSAD 5.x, Eclipse Galileo/Europa/3.x/2.x, MyEclipse 3.x/2.x, NetBeans 7.x/6.x, IntelliJ 7.x, Workshop 8.1/6.1, Adobe Photoshop, Adobe Dreamweaver, Adobe Flash, Ant, Maven, Rational Rose, RSA, MS Visio
RDBMS: SQL Server, Teradata, Oracle, MySQL
Languages: C, C++, Java, Python, OpenGL, APEX, VB .NET, XML, MATLAB, Microsoft Access, PostgreSQL, Perl, Java Script, Linux Bash Shell Scripting
Web Technologies: HTML5, CSS3, JavaScript, PHP, Ruby on Rails, ASP, ASP.NET JSP, JSF, Servlets, EJB, JDBC, Struts, Spring, Spring MVC, Spring Portlet, Spring Web Flow, Hibernate, iBATIS, JMS, MQ, JCA, JNDI, Java Beans, JAX-RPC, JAX-WS, RMI, RMI-IIOP, EAD4J, Axis, Castor, SOAP, WSDL, UDDI, JiBX, JAXB, DOM, SAX, MyFaces (Tomahawk), Facelets, JPA, Portal, Portlet, JSR 168/286, LifeRay, WebLogic Portal, LDAP, Junit.NET
Operating Systems: Windows, UNIX, Linux, Mac OS X Tools: Jenkins, Maven, GitHub, Tableau, Erwin Data Modeler, Weka, Informatica, Excel, Netbeans IDE, Eclipse
PROFESSIONAL EXPERIENCE:
Confidential, Round Rock, TX
Hadoop / Spark Developer
Responsibilities:
- Involved in all phases of Software Development Life Cycle (SDLC) activities such as development, implementation and support for Hadoop
- Imported and exported data from different Relational Databases like Mysql and Oracle into HDFS and Hive and vice-versa, using Sqoop
- Developed data pipe line using Kafka, HBase, Spark and Hive to ingest, transform and analyze data
- Migrated complex Map Reduce programs into Apache Spark RDD transformations
- Used Spark Streaming to divide streaming data into batches as an input to Spark engine for batch processing
- Developed Spark code and Spark-SQL/Streaming for faster testing and processing of data and handled Data Skewness in Spark-SQL
- Imported and exported data from different Relational Databases like Mysql and Oracle into HDFS and Hive and vice-versa, using Sqoop
- Developed data pipe line using Kafka, HBase, Spark and Hive to ingest, transform and analyze data
- Migrated complex Map Reduce programs into Apache Spark RDD transformations
- Created HBase column families to store various data types coming from various sources.
- Convert Hive/SQL queries into Spark transformations using Spark RDDs, Scala.
- Implemented Data Lakein HiveDB providing centralized platform for users to rapidly query, join, analyze from multiple sources at one point.
- Developed MapReduce programs to parse the raw data, populate staging tables and store the refined data in partitioned tables in the EDW.
- Have handled varying data formats like JSON, XML, ORC, parquet, sequence files etc.
- Worked with several Python libraries like NumPy, Pandas.
- Installed and configured MapReduce, HIVE and the HDFS; implemented CDH3 Hadoop cluster on CentOS. Assisted with performance tuning and monitoring.
- Have implemented free text search engines using Apache Solr collections.
- Have implemented end to end data transformation process in Oozie including data collection, decoding, transform, de-dup and storage using Hadoop frameworks like Map reduce, Hive, Pig.
- Experience in ingesting, deduping and processing machine generated raw data with Spark frameworks on Scala.
- Create Hive tables and date load/write Hive UDFs.
- Load and transform large sets of structured, semi-structured and unstructured data and analyzed them by running Hive queries and Pig scripts.
- Manage AWS EC2 instances utilizing Auto Scaling, Elastic Load Balancing and Glacier for our QA and UAT environments as well as infrastructure servers for GIT and Puppet
- Used the Spark - Cassandra Connector to load data to and from Cassandra.
- Real time streaming of data using Spark with Kafka.
- Configured Spark streaming to receive real time data from the Kafka and store the stream data to HDFS and store in databases such as HBase.
- Performed advanced procedures like text analytics and processing, using the in-memory computing of Spark using Scala
- Performed Continuous Integration-Continuous Deployment(CICD) using Jenkins and Maven
- Developed complex queries and User Defined Functions to extend core functionality of PIG & HIVE for data analysis
- Migrate the entire Data Centers to the cloud using VPC, EC2, S3, EMR, RDS, Splice Machine and DynamoDB services
- Scheduled workflow using Oozie for Map Reduce jobs, Pig & Hive Queries and managed cluster resources using Zookeeper
Environment: Hadoop, HDFS, Pig, Hive, Sqoop, Kafka, Zookeeper, Spark, Python, Hbase, Scala, Shell Scripting, Maven, MapReduce, Amazon EMR, EC2, S3
Confidential, Woonsocket, RI
Hadoop Developer
Responsibilities:
- Imported data using Sqoop to load data from MySQL to HDFS on regular basis
- Involved in collecting and aggregating large amounts of log data using Apache Flume and staging data in HDFS for further analysis
- Collected and aggregated large amounts of web log data from different sources such as webservers, mobile and network devices using Apache Kafka and stored the data into HDFS for analysis
- Developed multiple Kafka Producers and Consumers from scratch implementing organization's requirements
- Responsible for creating, modifying topics (Kafka Queues) as and when required with varying configurations involving replication factors, partitions and TTL
- Wrote and tested complex MapReduce jobs for aggregating identified and validated data
- Created Managed and External Hive tables with static/dynamic partitioning
- Extensively involved in performance tuning of the HiveQLby performing bucketing on large Hive tables
- Implemented Spark applications using Scala and Spark SQL for faster testing and processing of data
- Developed an equivalent Spark Scala code for existing SAS code to extract summary insights on the Hive tables
- Integrated Amazon Redshift with Spark using Scala
- Extensive experience in writing Pig scripts to transform raw data from several data sources into forming baseline data
- Implemented workflow using Oozie for running Map Reduce jobs and Hive Queries
- Expert in implementing advanced procedures like text analytics and processing using the in-memory computing capabilities like Apache Spark written in Scala.
- Developed and executed shell scripts to automate the jobs.
- Wrote complex Hive queries and UDFs.
- Worked on reading multiple data formats on HDFS using PySpark.
- Involved in converting Hive/SQL queries into Spark transformations using Spark RDDs, Python and Scala.
- Developed multiple POCs using PySpark and deployed on the Yarn cluster, compared the performance of Spark, with Hive and SQL/Teradata.
- Analyzed the SQL scripts and designed the solution to implement using PySpark.
- Involved in loading data from UNIX file system to HDFS.
- Extracted the data from Teradata into HDFS using Sqoop.
- Handled importing of data from various data sources, performed transformations using Hive, Map Reduce, Spark and loaded data into HDFS.
- Manage and review Hadoop log files.
- Involved in analysis, design, testing phases and responsible for documenting technical specifications.
- Developed Kafka producer and consumers, HBase clients, Spark and Hadoop MapReduce jobs along with components on HDFS, Hive.
- Very good understanding of Partitions, Bucketing concepts in Hive and designed both Managed and External tables in Hive to optimize performance.
- Worked on the core and Spark SQL modules of Spark extensively.
- Experienced in running Hadoop streaming jobs to process terabytes data.
- Involved in importing the real-time data to Hadoop using Kafka and implemented the Oozie job for daily imports.
Environment: Hadoop, HDFS, Map Reduce, Hive, Sqoop, Apache Kafka, Zookeeper, Spark, Hbase, Python, Shell Scripting, Oozie, Maven, Hortonworks.
Confidential, New York City
Hadoop Developer
Responsibilities:
- Responsible for building scalable distributed data solutions using Hadoop
- Worked comprehensively with Apache Sqoop and developed Sqoop scripts to interface data from a MySQL database into the Hadoop Distributed File System (HDFS)
- Utilize parallel processes of the Hadoop Framework to ensure resource efficiency
- Created Managed tables and External tables in Hive and loaded data from HDFS
- Worked on debugging and performance tuning of Hive & Pig Jobs
- Used python sub-process module to perform UNIX shell commands
- Extracted data from Agent Nodes into HDFS using Python scripts
- Implemented HiveGeneric UDF's to in corporate business logic into Hive Queries
- Analysed the web log data using the HiveQL to extract number of unique visitors per day, page views, visit duration, most visited page on website etc.
- Used pig and hive Upon the Hcatalog tables to analyse the data and create schema for the Hbase tables in Hive
- Coordinated with the BI team to visualize the transformed data into a dashboard using Tableau
- Assisted in creating and maintaining Technical documentation to launching Hadoop Clusters and executing Hive queries and Pig Scripts.
- Created HBase tables to load large sets of structured, semi-structured and unstructured data coming from UNIX, NoSQL and a variety of portfolios.
- Supported code/design analysis, strategy development and project planning.
- Created reports for the BI team usingSqoop to export data into HDFS and Hive.
- Developed multiple MapReduce jobs in Java for data cleaning and preprocessing.
- Assisted with data capacity planning and nodeforecasting.
- Collaborated with the infrastructure, network, database, application and BI teams to ensure data quality and availability.
- Administrator for Pig, Hive and Hbase installing updates, patches and upgrades
Environment: Hadoop, HDFS, Map Reduce, Hive, Sqoop, Zookeeper, Hbase, Python, Shell Scripting, Oozie, Maven, Cloudera, Tableau
Confidential
Java Developer
Responsibilities:
- Involved in various phases of Software Development Life Cycle (SDLC) of the application like Requirements gathering, Design, Analysis and Code development and deployment into various environments.
- Developed user interface layer using Spring framework.
- Created new RESTweb service operations and modified the existing web service'sWADLs Web Application Description Language
- Involved in consuming, producing SOAP based web services using JAX-WS.
- Integrated the ORMObject Relational Mapping tool hibernate to the spring using Spring ORM in our app and used spring transaction API for database related transactions.
- Implemented Data Access Object core java pattern to abstract and encapsulate all access to the data source and singleton pattern.
- Involved in back end development using Oracle, triggers, views and stored procedure (PL/SQL).
- Involved in batch processing using Spring Batch framework to extract data from database and load into corresponding Credit card tables
- Led the migration of monthly statements from UNIX platform to MVC Web-based Windows application using Java, JSP, and Struts technology.
- Prepared use cases, designed and developed object models and class diagrams.
- Developed SQL statements to improve back-end communications.
- Incorporated custom logging mechanism for tracing errors, resolving all issues and bugs before deploying the application in the WebSphere Server.
- Received praise from users, shareholders and analysts for developing a highly interactive and intuitive UI using JSP, AJAX, JSF and JQuery techniques.
- Used SOAP UI tool to test the REST/SOAP web service operations
- Used Atlassian Jira for tracking the user stories in agile methodology in development.
- Responsible for Unit and UAT testing of the developed part of applications.
- Debugging and fixing of any developmental/environmental issues
Environment: J2EE, SOAP/REST Web Services, Spring, Hibernate, Spring Batch, SOAP UI, Ant, Gradle, JMeter, WSDL, SVN, JAXB, XML, JDBC, HTML, CSS, and Eclipse
Confidential
Java Developer
Responsibilities:
- Involved in various phases of Software Development Life Cycle, such as requirements gathering, modelling, analysis, design and development
- Ensured clear understanding of customer's requirements before developing the final proposal
- Generated Use case diagrams, Activity flow diagrams, Class diagrams and Object diagrams in the design phase
- Used Java Design Patterns like DAO, Singleton etc
- Written complex SQL queries for retrieving and updating data
- Involved in implementing multithreaded environment to generate messages
- Used JDBC Connections and WebSphere Connection pool for database access
- Used Struts tag libraries (like html, logic, tab, bean etc.) and JSTL tags in the JSP pages
- Involved in development using Struts components - Struts-config.xml, tiles, form-beans and plug-ins in Struts architecture
- Involved in design and implementation of document based Web Services
- Used prepared statements and callable statements to implement batch insertions and access stored procedures
- Involved in bug fixing and for the new enhancements
- Responsible for handling the production issues and provided solutions
- Configured connection pooling using WebLogic application server
- Developed and Deployed the Application on WebLogic using ANT build.xml script
- Developed SQL queries and stored procedures to execute the backend processes using Oracle
- Deployed application on WebLogic Application Server and development using Eclipse.
Environment: Java 1.4, Servlets, JSP, JMS, Struts, Validation Framework, tag Libraries, JSTL, JDBC, PL/SQL, HTML, JavaScript, Oracle 9i (SQL), UNIX, AJAX, Eclipse 3.0, LINUX, CV
