We provide IT Staff Augmentation Services!

Hadoop / Spark Developer Resume

3.00/5 (Submit Your Rating)

Round Rock, TX

PROFESSIONAL SUMMARY:

  • 8+ years of IT experience in Analysis, Design, Development, Implementation, Integration and testing of Application Software in web - based environments, distributed n-tier products and Client/Server architectures.
  • 5+ years of Big Data/Hadoop experience using Hadoop stack (MapReduce Programming, Pig, Hive, Oozie and Sqoop, Flume, Spark/Spark SQL, Strom, Kafka, Yarn/MRv2) and NoSQL technologies such as Cassandra, Hbase and MongoDB.
  • Good knowledge of Hadoop Architecture and various components such as HDFS, Job Tracker, Task Tracker, Name Node, Data Node, Resource manager and YARN concepts.
  • Sophisticated experience with Hadoop distributions like Apache, IBM Big Insights, Cloudera, Hortonworks&MapR.
  • Expert in importing data using Sqoop into HDFS from various Relational Database Systems.
  • Worked on standards and proof of concept in support of CDH4 and CDH5 implementation using AWScloud infrastructure.
  • Hands on experience in implementing complex business logic and optimizing the query using Hive QL and controlling the data distribution by partitioning and bucketing techniques to enhance performance.
  • Hands on experience in working with ClouderaCDH3 and CDH4 platforms
  • Good knowledge in Architecting, Designing, re-Engineering and Performance Optimization.
  • Experience with Oozie Workflow Engine in running workflow jobs with actions that run Hadoop Map Reduce and Pig jobs.
  • Used Talend Open Studio for Big Data Integration and ETL operations.
  • Experienced in creating and analyzing Software Requirement Specifications (SRS) and Functional Specification Document (FSD).
  • Worked extensively with Dimensional modeling, Data migration, Data cleansing, Data profiling, and ETL Processes features for data warehouses.
  • Expertise in workflow management tools like SQL Workbench, SQL Developer and TOAD tool for accessing the Database server.
  • Very Good Knowledge in Logicaland Physical Data modeling, creating new data models, data flows and data dictionaries.
  • Experience in understanding the security requirements for Hadoop and integrate with Kerberos authentication and authorization infrastructure.
  • Experience in importing and exporting the data using Sqoop from HDFS to Relational Database systems/mainframe and vice-versa.
  • Hands-on experience in using relational databases like Oracle, MySQL and MS-SQL Server.
  • Experienced the integration of various data sources like Java, RDBMS, Shell Scripting, Spreadsheets, and Text files.
  • Experience in Web Services using XML, HTML and SOAP.
  • Excellent Java development skills using Java 6/7/8, J2EE, J2SE, Servlets, Junit, JSP, JDBC.
  • Experience using middleware architecture using Sun Java technologies like J2EE, JSP, Servlets, and application servers like WebSphere and Web logic.
  • Expertise in creating tables and distributing data in Hive by implementing Partitioning and Bucketing, writing complex scripts and User Defined Functions in Pig & Hive and customized MapReduce jobs in Java.
  • Experience in working with Spark tools like RDD transformations, Spark MLlib and spark SQL
  • Experiencein working with NoSQL database HBase, MongoDB and Cassandra
  • Specialized in implementing Spark and Scala applications using higher order functions for both batch and interactive analysis.
  • Experience in working withAmazon AWS services such as EMR, EC2, S3, etc.
  • Worked and learned a great deal from Amazon Web Services (AWS) Cloud services like EC2, S3, EBS, RDS and VPC.
  • Ability to spin up different AWS instances including EC2-classic and EC2-VPC using cloud formation templates.
  • Expertise in using IDE like Eclipse, WebSphere (WSAD), NetBeans, My Eclipse, WebLogic Workshop
  • Experience in developing and designing Web Services (SOAP and Restful Web services)
  • Experience in developing Web Interface using Servlets, JSP and Custom Tag Libraries
  • Experience in developing applications using SCRUM methodology and Agile Methodology
  • Familiarity working with popular frameworks likes Struts, Hibernate, Spring MVC and AJAX.
  • Exceptional ability to learn and master new technologies and to deliver outputs in short deadlines.
  • Quick Learner, with high degree of passion and commitment in work.
  • Proficiency with mentoring and on-boarding new engineers who are not proficient in Hadoop and getting them up to speed quickly.
  • A good team player with excellent time management skills having profound insight to determine priorities, schedule work and meet critical deadlines.

TECHNICAL SKILLS:

Hadoop/Big Data: Apache Hadoop, MapReduce, Pig, Hive, Sqoop, Oozie, Flume, Zookeeper, Impala, Spark, Scala, Ambari, Impala, Kafka, YARN, HDFS, Ranger, Mapreduce, HBase, SOLR, Flume, MongoDB, Puppet, Oozie, Zookeeper, Spark, Talend, Hortonworks & Cloudera distributions

NOSQL Databases: HBase, MongoDB, Cassandra

Development Framework/IDE: RAD 8.x/7.x/6.0, IBM WebSphere Integration Developer 6.1, WSAD 5.x, Eclipse Galileo/Europa/3.x/2.x, MyEclipse 3.x/2.x, NetBeans 7.x/6.x, IntelliJ 7.x, Workshop 8.1/6.1, Adobe Photoshop, Adobe Dreamweaver, Adobe Flash, Ant, Maven, Rational Rose, RSA, MS Visio

RDBMS: SQL Server, Teradata, Oracle, MySQL

Languages: C, C++, Java, Python, OpenGL, APEX, VB .NET, XML, MATLAB, Microsoft Access, PostgreSQL, Perl, Java Script, Linux Bash Shell Scripting

Web Technologies: HTML5, CSS3, JavaScript, PHP, Ruby on Rails, ASP, ASP.NET JSP, JSF, Servlets, EJB, JDBC, Struts, Spring, Spring MVC, Spring Portlet, Spring Web Flow, Hibernate, iBATIS, JMS, MQ, JCA, JNDI, Java Beans, JAX-RPC, JAX-WS, RMI, RMI-IIOP, EAD4J, Axis, Castor, SOAP, WSDL, UDDI, JiBX, JAXB, DOM, SAX, MyFaces (Tomahawk), Facelets, JPA, Portal, Portlet, JSR 168/286, LifeRay, WebLogic Portal, LDAP, Junit.NET

Operating Systems: Windows, UNIX, Linux, Mac OS X Tools: Jenkins, Maven, GitHub, Tableau, Erwin Data Modeler, Weka, Informatica, Excel, Netbeans IDE, Eclipse

PROFESSIONAL EXPERIENCE:

Confidential, Round Rock, TX

Hadoop / Spark Developer

Responsibilities:

  • Involved in all phases of Software Development Life Cycle (SDLC) activities such as development, implementation and support for Hadoop
  • Imported and exported data from different Relational Databases like Mysql and Oracle into HDFS and Hive and vice-versa, using Sqoop
  • Developed data pipe line using Kafka, HBase, Spark and Hive to ingest, transform and analyze data
  • Migrated complex Map Reduce programs into Apache Spark RDD transformations
  • Used Spark Streaming to divide streaming data into batches as an input to Spark engine for batch processing
  • Developed Spark code and Spark-SQL/Streaming for faster testing and processing of data and handled Data Skewness in Spark-SQL
  • Imported and exported data from different Relational Databases like Mysql and Oracle into HDFS and Hive and vice-versa, using Sqoop
  • Developed data pipe line using Kafka, HBase, Spark and Hive to ingest, transform and analyze data
  • Migrated complex Map Reduce programs into Apache Spark RDD transformations
  • Created HBase column families to store various data types coming from various sources.
  • Convert Hive/SQL queries into Spark transformations using Spark RDDs, Scala.
  • Implemented Data Lakein HiveDB providing centralized platform for users to rapidly query, join, analyze from multiple sources at one point.
  • Developed MapReduce programs to parse the raw data, populate staging tables and store the refined data in partitioned tables in the EDW.
  • Have handled varying data formats like JSON, XML, ORC, parquet, sequence files etc.
  • Worked with several Python libraries like NumPy, Pandas.
  • Installed and configured MapReduce, HIVE and the HDFS; implemented CDH3 Hadoop cluster on CentOS. Assisted with performance tuning and monitoring.
  • Have implemented free text search engines using Apache Solr collections.
  • Have implemented end to end data transformation process in Oozie including data collection, decoding, transform, de-dup and storage using Hadoop frameworks like Map reduce, Hive, Pig.
  • Experience in ingesting, deduping and processing machine generated raw data with Spark frameworks on Scala.
  • Create Hive tables and date load/write Hive UDFs.
  • Load and transform large sets of structured, semi-structured and unstructured data and analyzed them by running Hive queries and Pig scripts.
  • Manage AWS EC2 instances utilizing Auto Scaling, Elastic Load Balancing and Glacier for our QA and UAT environments as well as infrastructure servers for GIT and Puppet
  • Used the Spark - Cassandra Connector to load data to and from Cassandra.
  • Real time streaming of data using Spark with Kafka.
  • Configured Spark streaming to receive real time data from the Kafka and store the stream data to HDFS and store in databases such as HBase.
  • Performed advanced procedures like text analytics and processing, using the in-memory computing of Spark using Scala
  • Performed Continuous Integration-Continuous Deployment(CICD) using Jenkins and Maven
  • Developed complex queries and User Defined Functions to extend core functionality of PIG & HIVE for data analysis
  • Migrate the entire Data Centers to the cloud using VPC, EC2, S3, EMR, RDS, Splice Machine and DynamoDB services
  • Scheduled workflow using Oozie for Map Reduce jobs, Pig & Hive Queries and managed cluster resources using Zookeeper

Environment: Hadoop, HDFS, Pig, Hive, Sqoop, Kafka, Zookeeper, Spark, Python, Hbase, Scala, Shell Scripting, Maven, MapReduce, Amazon EMR, EC2, S3

Confidential, Woonsocket, RI

Hadoop Developer

Responsibilities:

  • Imported data using Sqoop to load data from MySQL to HDFS on regular basis
  • Involved in collecting and aggregating large amounts of log data using Apache Flume and staging data in HDFS for further analysis
  • Collected and aggregated large amounts of web log data from different sources such as webservers, mobile and network devices using Apache Kafka and stored the data into HDFS for analysis
  • Developed multiple Kafka Producers and Consumers from scratch implementing organization's requirements
  • Responsible for creating, modifying topics (Kafka Queues) as and when required with varying configurations involving replication factors, partitions and TTL
  • Wrote and tested complex MapReduce jobs for aggregating identified and validated data
  • Created Managed and External Hive tables with static/dynamic partitioning
  • Extensively involved in performance tuning of the HiveQLby performing bucketing on large Hive tables
  • Implemented Spark applications using Scala and Spark SQL for faster testing and processing of data
  • Developed an equivalent Spark Scala code for existing SAS code to extract summary insights on the Hive tables
  • Integrated Amazon Redshift with Spark using Scala
  • Extensive experience in writing Pig scripts to transform raw data from several data sources into forming baseline data
  • Implemented workflow using Oozie for running Map Reduce jobs and Hive Queries
  • Expert in implementing advanced procedures like text analytics and processing using the in-memory computing capabilities like Apache Spark written in Scala.
  • Developed and executed shell scripts to automate the jobs.
  • Wrote complex Hive queries and UDFs.
  • Worked on reading multiple data formats on HDFS using PySpark.
  • Involved in converting Hive/SQL queries into Spark transformations using Spark RDDs, Python and Scala.
  • Developed multiple POCs using PySpark and deployed on the Yarn cluster, compared the performance of Spark, with Hive and SQL/Teradata.
  • Analyzed the SQL scripts and designed the solution to implement using PySpark.
  • Involved in loading data from UNIX file system to HDFS.
  • Extracted the data from Teradata into HDFS using Sqoop.
  • Handled importing of data from various data sources, performed transformations using Hive, Map Reduce, Spark and loaded data into HDFS.
  • Manage and review Hadoop log files.
  • Involved in analysis, design, testing phases and responsible for documenting technical specifications.
  • Developed Kafka producer and consumers, HBase clients, Spark and Hadoop MapReduce jobs along with components on HDFS, Hive.
  • Very good understanding of Partitions, Bucketing concepts in Hive and designed both Managed and External tables in Hive to optimize performance.
  • Worked on the core and Spark SQL modules of Spark extensively.
  • Experienced in running Hadoop streaming jobs to process terabytes data.
  • Involved in importing the real-time data to Hadoop using Kafka and implemented the Oozie job for daily imports.

Environment: Hadoop, HDFS, Map Reduce, Hive, Sqoop, Apache Kafka, Zookeeper, Spark, Hbase, Python, Shell Scripting, Oozie, Maven, Hortonworks.

Confidential, New York City

Hadoop Developer

Responsibilities:

  • Responsible for building scalable distributed data solutions using Hadoop
  • Worked comprehensively with Apache Sqoop and developed Sqoop scripts to interface data from a MySQL database into the Hadoop Distributed File System (HDFS)
  • Utilize parallel processes of the Hadoop Framework to ensure resource efficiency
  • Created Managed tables and External tables in Hive and loaded data from HDFS
  • Worked on debugging and performance tuning of Hive & Pig Jobs
  • Used python sub-process module to perform UNIX shell commands
  • Extracted data from Agent Nodes into HDFS using Python scripts
  • Implemented HiveGeneric UDF's to in corporate business logic into Hive Queries
  • Analysed the web log data using the HiveQL to extract number of unique visitors per day, page views, visit duration, most visited page on website etc.
  • Used pig and hive Upon the Hcatalog tables to analyse the data and create schema for the Hbase tables in Hive
  • Coordinated with the BI team to visualize the transformed data into a dashboard using Tableau
  • Assisted in creating and maintaining Technical documentation to launching Hadoop Clusters and executing Hive queries and Pig Scripts.
  • Created HBase tables to load large sets of structured, semi-structured and unstructured data coming from UNIX, NoSQL and a variety of portfolios.
  • Supported code/design analysis, strategy development and project planning.
  • Created reports for the BI team usingSqoop to export data into HDFS and Hive.
  • Developed multiple MapReduce jobs in Java for data cleaning and preprocessing.
  • Assisted with data capacity planning and nodeforecasting.
  • Collaborated with the infrastructure, network, database, application and BI teams to ensure data quality and availability.
  • Administrator for Pig, Hive and Hbase installing updates, patches and upgrades

Environment: Hadoop, HDFS, Map Reduce, Hive, Sqoop, Zookeeper, Hbase, Python, Shell Scripting, Oozie, Maven, Cloudera, Tableau

Confidential

Java Developer

Responsibilities:

  • Involved in various phases of Software Development Life Cycle (SDLC) of the application like Requirements gathering, Design, Analysis and Code development and deployment into various environments.
  • Developed user interface layer using Spring framework.
  • Created new RESTweb service operations and modified the existing web service'sWADLs Web Application Description Language
  • Involved in consuming, producing SOAP based web services using JAX-WS.
  • Integrated the ORMObject Relational Mapping tool hibernate to the spring using Spring ORM in our app and used spring transaction API for database related transactions.
  • Implemented Data Access Object core java pattern to abstract and encapsulate all access to the data source and singleton pattern.
  • Involved in back end development using Oracle, triggers, views and stored procedure (PL/SQL).
  • Involved in batch processing using Spring Batch framework to extract data from database and load into corresponding Credit card tables
  • Led the migration of monthly statements from UNIX platform to MVC Web-based Windows application using Java, JSP, and Struts technology.
  • Prepared use cases, designed and developed object models and class diagrams.
  • Developed SQL statements to improve back-end communications.
  • Incorporated custom logging mechanism for tracing errors, resolving all issues and bugs before deploying the application in the WebSphere Server.
  • Received praise from users, shareholders and analysts for developing a highly interactive and intuitive UI using JSP, AJAX, JSF and JQuery techniques.
  • Used SOAP UI tool to test the REST/SOAP web service operations
  • Used Atlassian Jira for tracking the user stories in agile methodology in development.
  • Responsible for Unit and UAT testing of the developed part of applications.
  • Debugging and fixing of any developmental/environmental issues

Environment: J2EE, SOAP/REST Web Services, Spring, Hibernate, Spring Batch, SOAP UI, Ant, Gradle, JMeter, WSDL, SVN, JAXB, XML, JDBC, HTML, CSS, and Eclipse

Confidential

Java Developer

Responsibilities:

  • Involved in various phases of Software Development Life Cycle, such as requirements gathering, modelling, analysis, design and development
  • Ensured clear understanding of customer's requirements before developing the final proposal
  • Generated Use case diagrams, Activity flow diagrams, Class diagrams and Object diagrams in the design phase
  • Used Java Design Patterns like DAO, Singleton etc
  • Written complex SQL queries for retrieving and updating data
  • Involved in implementing multithreaded environment to generate messages
  • Used JDBC Connections and WebSphere Connection pool for database access
  • Used Struts tag libraries (like html, logic, tab, bean etc.) and JSTL tags in the JSP pages
  • Involved in development using Struts components - Struts-config.xml, tiles, form-beans and plug-ins in Struts architecture
  • Involved in design and implementation of document based Web Services
  • Used prepared statements and callable statements to implement batch insertions and access stored procedures
  • Involved in bug fixing and for the new enhancements
  • Responsible for handling the production issues and provided solutions
  • Configured connection pooling using WebLogic application server
  • Developed and Deployed the Application on WebLogic using ANT build.xml script
  • Developed SQL queries and stored procedures to execute the backend processes using Oracle
  • Deployed application on WebLogic Application Server and development using Eclipse.

Environment: Java 1.4, Servlets, JSP, JMS, Struts, Validation Framework, tag Libraries, JSTL, JDBC, PL/SQL, HTML, JavaScript, Oracle 9i (SQL), UNIX, AJAX, Eclipse 3.0, LINUX, CV

We'd love your feedback!