Hadoop-spark Developer Resume
Philadelphia, PA
SUMMARY:
- Having 8+ years of experience in software design, development, implementation and support of various applications like Big Data (Hadoop) and Java technologies.
- 3.6 years of experience with Hadoop Ecosystem including Spark, Scala,HDFS, Map Reduce, Hive, Pig, Storm, Kafka, YARN, HBase, Oozie, Zookeeper, Flume, Sqoop,
- Assisted in Cluster maintenance, Cluster Monitoring and Troubleshooting, Managing and Reviewing data backups and log files.
- Excellent ability to use analytical tools to mine data, Predictive analysis, evaluating the underlying patterns and implement complex algorithms for data analysis.
- 1.5 Year Hands On experience on SPARK, Spark Streaming, Spark Mlib, SCALA,
- Creating the Data Frames handle in SPARK with Scala.
- Hands On experience on developing UDF, DATA Frames and SQL Queries in SPARK SQL
- Developed PIG Latin scripts and SPARKSQL scripts for handling data formation.
- Hands on experience on Real Time data tools like Kafka and Storm.
- Developed SQOOP Scripts for importing large dataset from RDBMS to HDFS
- Experience in writing & Creating the UDF’s in Java and Register them in PIG and HIVE
- Good understanding on Spark architecture and its components.
- Experience in writing Pig Latin Scripts.
- Efficient in writing the Map Reduce programs for analyzing structured and unstructured data.
- Expertise in working with Hive data warehouse tool - creating tables, data distribution by implementing partitioning and bucketing, writing and optimizing the HiveQL queries.
- Experience in using Apache Sqoop to import and export data to and from HDFS and Hive.
- Hands on experience in setting up workflow using Apache Oozie workflow engine for managing and scheduling Hadoop job
- Experience in scheduling the jobs using Oozie Coordinator, Bundler and Crontab. Cloud Infrastructure:
- Experience with AWS components like Amazon Ec2 instances, S3 buckets and Cloud Formation templates and Boto library.
- Experience with Azure Components like Azure SQl Database and Data Factory. File Formats:
- Experienced in working with different file formats - Avro, Parquet, RC and ORC.
- Hands on experience in configuring and working with Flume to load the data from multiple sources directly into HDFS.
- Experienced and skilled Agile Developer with a strong record of excellent teamwork and Successful coding.
TECHNICAL SKILLS:
Hadoop Technologies and Distributions: Apache Hadoop, Cloudera Hadoop Distribution CDH3, CDH4, CDH5 and Horton works Data Platform (HDP)
Hadoop Ecosystem: HDFS, Hive, Pig, Sqoop, Oozie, Flume, Spark, Zookeeper, Map-Reduce, Spark-SQL, Spark Streaming and Spark MLib.
NoSQL Databases: HBase, Cassandra
Programming: C, C++,Python, Java, SCALA,PL/SQL,SBT,MAVEN
RDBMS: ORACLE, MySQL, SQL Server
Web Development: HTML, JSP, Servlets, JavaScript, CSS, XML
IDE: Eclipse4.x, NetBeans, Microsoft Visual Studio
Operating Systems: Linux (RedHat, CentOS), Windows XP/7/8 and Z/OS(Main Frames)
Web Servers: Apache Tomcat
Cluster Management Tools: Cloudera Manager, Horton Works Ambari and Hadoop Security Tools
PROFESSIONAL EXPERIENCE:
Confidential, Philadelphia, PA
Hadoop-Spark Developer
Responsibilities:- Importing data from relational data stores to Hadoop using Sqoop.
- Creating various MapReduce jobs for performing ETL transformations on the transactional and application specific data sources.
- Wrote and executed PIG scripts using Grunt shell.
- Big data analysis using Pig and User defined functions (UDF).
- Performed joins, group by and other operations in MapReduce by using Java and PIG.
- Processed and formatted the output from PIG, Hive before sending to the Hadoop output file.
- Used HIVE definition to map the output file to tables.
- Wrote map reduce/HBase jobs.
- Reviewed the HDFS usage and system design for future scalability and fault-tolerance.
- Worked with HBASE NOSQL database.
- Experienced in analyzing Cassandra database and compare it with other open-source NoSQL databases to find which one of them better suites the current requirements.
- Created UDF’s to encrypt the customer sensitive data and stored into HDFS and performed analysis using PIG.
- Developed spark scripts by using Scala shell as per requirements.
- Developed multiple POCs using Spark and deployed on the Yarn cluster, compared the performance of Spark, with Hive and SQL.
- Developed Kafka producer and consumers, Cassandra clients and Spark along with components on HDFS, Hive.
- Involved in converting Hive/SQL queries into Spark transformations using Spark RDDs, Python and Scala.
- Used HUE for running Hive queries. Created partitions according to day using Hive to improve performance.
Environment: Cloudera CDH 5.12 Hadoop 2.6.0, MapReduce, Hive, HBase, HDFS, Cassandra, PIG, Nifi, Sqoop, Oozie, Java 1.8, UNIX, Python, Scala and Spark.
Confidential, Irving, TX
Hadoop-Spark Developer
Responsibilities:- Involved in creating Hive tables, and loading and analyzing data using hive queries.
- Optimizing of existing algorithms in Hadoop using Spark Context, Spark-Sql, Data Frames and Pair RDD’s.
- Worked on Cluster of size 135 nodes.
- Implemented Spark using Scala and utilizing Data frames and Spark SQL API for faster processing of data
- Created RDD’s, Data Frames and Datasets.
- Created Hive Tables, loaded data from Teradata using Sqoop.
- Worked on tuning of back-end stored procedures using TOAD.
- Good experience with Talend Open Studio for designing ETL Jobs for Processing of data.
- Used ORC, Parquet file formats for storing the data.
- Used java code for Sql Queries and also code to retrieve the Sql Queries through Text File.
- Responsible for troubleshooting issues in the execution of MapReduce jobs by inspecting and reviewing log files.
- Used Sqoop to transfer data between RDBMS and Hadoop Distributed File System.
- Used Eclipse for the Development, Testing and Debugging of the application.
- Use python for writing script to move the data cluster to cluster.
- Log4j framework has been used for logging debug, info & error data.
- Created Hive External and Managed tables.
- Use of MAVEN for dependency management and structure of the project.
- Designed and Maintained Tez workflows to manage the flow of jobs in the cluster
- Loaded the Spark RDD and do in memory data Computation to generate the Output response.
- Sometimes variable needs to share across the Nodes. So in such cases we used shared variables: Broadcast Variable, Accumulator.
Environment: Hadoop, MapReduce, Hive, pig, spring batch, Scala, Sqoop, Bash Scripting, Spark RDD, Spark Sql, Spark Data Frames, Maven Eclipse.
Confidential,Des Moines, IA
Hadoop Developer
Responsibilities:- Involved in Design and Development of technical specifications.
- Written shell scripts to pull the data from Tumbleweed server to cornerstone staging area.
- Data conversion from EBICIDC to ASCII format.
- Written Sqoop commands to pull the data from Teradata Source.
- Written Pig scripts to preprocess the data before loading to cornerstone.
- Optimization of Hive Scripts.
- Registration of feeds metadata in MYSQL tables.
- Written shell scripts and scheduled our jobs through UNIX crons.
- Written Job Workflows using Spring Batch.
- Worked on Project deployment from Gold cluster to platinum Cluster.
- Provide support for PRD Support Team.
- Closely worked with Hadoop security team and infrastructure team to implement security.
- Implemented authentication and authorization service using Kerberos authentication Protocol.
- Designed and implemented streaming data on UI with Scala.js
- Hands-on experience with systems-building languages such as Scala, Java
- Programs for Validation/Normalizing/Enriching and REST API to Develop UI Based on manual QA Validation. Used SparkSQL, Scala to running QA based SQL queries.
- Creating RDD's and Pair RDD's for Spark Programming.
- Implement Joins, Grouping and Aggregations for the Pair RDD's.
- Save the result in HIVE for the downstream to access the data.
- Use Dataframes for data transformations.
Environment: Hadoop, MapReduce, Hive, pig, spring batch, Scala, Sqoop, Bash Scripting, Spark RDD, Spark Sql.
Confidential,Santa Ana, CA
Hadoop Developer
Responsibilities:- Developed Big Data Solutions that enabled the business and technology teams to make data-driven decisions on the best ways to acquire customers and provide them business solutions.
- Involved in installing, configuring and managing Hadoop Ecosystem components like Spark, Hive, Pig, Sqoop, Kafka and Flume.
- Involved in installing Hadoop and Spark Cluster in Amazon Web Server.
- Work Amazon Ec2 instances, S3 buckets and Cloud Formation templates and Boto library.
- Migrated the existing data to Hadoop from RDBMS (SQL Server and Oracle) using Sqoop for processing the data.
- Responsible for Data Ingestion like Flume and Kafka.
- Responsible for loading unstructured and semi-structured data into Hadoop cluster coming from different sources using Flume and managing.
- Developed Spark Programs for Batch and Real Time Processing.
- Developed Spark Streaming applications for Real Time Processing.
- Developed MapReduce programs to cleanse and parse data in HDFS obtained from various data sources and to perform joins on the Map side using distributed cache.
- Used Hive data warehouse tool to analyze the data in HDFS and developed Hive queries.
- Created internal and external tables with properly defined static and dynamic partitions for efficiency.
- Used the RegEx, JSON and Avro SerDe’s for serialization and de-serialization packaged with Hive to parse the contents of streamed log data.
- Implemented Hive custom UDF’s to achieve comprehensive data analysis.
- Uses Talend Open Studio to load files into Hadoop HIVE tables and performed ETL Aggregations in Hadoop HIVE.
- Designing & Creating ETL Jobs through Talend to load huge volumes of data into Cassandra, Hadoop Ecosystem and relational databases.
- Implemented authentication and authorization service using Kerberos authentication Protocol
- Used Pig to develop ad-hoc queries.
- Exported the business required information to RDBMS using Sqoop to make the data available for BI team to generate reports based on data.
- Implemented daily workflow for extraction, processing and analysis of data with Oozie.
- Responsible for troubleshooting MapReduce jobs by reviewing the log files.
Environment: Hadoop, Spark, Spark Streaming, Spark Mlib, Scala Hive, Pig, Hcatalog, MapReduce, Oozie, Sqoop, Flume and Kafka, Kerberos.
Confidential, Charlotte, NC
Sr. Java Developer
Responsibilities:- Analyzed and reviewed client requirements and design.
- Followed agile methodology for development process.
- Developed presentation layer using HTML5, and CSS3, Ajax.
- Developed the application using Struts Framework that uses Model View Controller (MVC)
- Architecture with JSP as the view.
- Involved in Performance Tuning
- Extensively used Spring IOC for Dependency Injection and worked on Custom MVC
- Frameworks loosely based on Struts.
- Used Restful Web services for transferring data between applications.
- Configured spring with ORM framework Hibernate for handling DAO classes and to bind Objects to the relational model.
- Developed Web services (JAX-WS) specification using Apache CXF as the implementation and developed client application API's using NodeJs.
- Adopted J2EE design patterns like Singleton, Service Locator and Business Facade.
- Developed POJO classes and used annotations to map with database tables.
- Used Java Message Service (JMS) for reliable and asynchronous exchange of important Information such as Credit card transactions report.
- Used Multi-Threading to handle more users.
- Developed Hibernate JDBC code for establishing communication with database.
- Worked with DB2 database for persistence with the help of PL/SQL querying.
- Used SQL queries to retrieve information from database.
- Developed various triggers, functions, procedures, views for payments.
- XSL/XSLT is used for transforming and displaying reports.
- Used GIT to keep track of all work and all changes in source code.
- Used JProfiler for performance tuning.
- Wrote test cases which adhere to a Test Driven Development (TDD) pattern.
- Used Junit, a test framework which uses annotations to identify methods that specify a Test.
- Used Log 4J to log messages depending on the messages type and level.
- Built the application using MAVEN and deployed using Web Sphere Application server.
Environment: Java 8, Spring framework, Spring Model View Controller (MVC), Struts 2.0, XML, Hibernate 3.0,UML, Java Server Pages (JSP) 2.0, Servlets 3.0, JDBC4.0, Junit, Log4j, MAVEN, Win 7, HTML, REST Client Eclipse, Agile Methodology, Design Patterns, Web Sphere 6.1.Java/J2EE Developer.
Confidential
Java Developer
Responsibilities:- Implemented Microsoft Visio and Rational Rose for designing the Use Case Diagrams, Class
- Model, Sequence diagrams, and Activity diagrams for SDLC process of the application.
- Deployed GUI pages by using JSP, JSTL, HTML, DHTML, XHTML, CSS, JavaScript, AJAX
- Configured the project on Web Sphere 6.1 application servers
- Implemented the online application by using Core Java, Jdbc, JSP, Servlets and EJB 1.1, Web Services, SOAP, WSDL
- Communicated with other Health Care info by using Web Services with the help of SOAP, WSDL JAX-RPC
- Used Singleton, factory design pattern, DAO Design Patterns based on the application requirements.
- Used SAX and DOM parsers to parse the raw XML documents
- Used RAD as Development IDE for web applications.
- Preparing and executing Unit test cases
- Used Log4J logging framework to write Log messages with various levels.
- Involved in fixing bugs and minor enhancements for the front-end modules.
- Doing functional and technical reviews
- Maintenance in the testing team for System testing/Integration/UAT
- Guaranteeing quality in the deliverables.
- Conducted Design reviews and Technical reviews with other project stakeholders.
- Was a part of the complete life cycle of the project from the requirements to the production support.
- Created test plan documents for all back end database modules.
- Implemented the project in Linux environment.
- Environment: JDK 1.5, JSP, Web Sphere, JDBC, EJB2.0, XML, DOM, SAX, XSLT, CSS, HTML, JNDI, Web.
Environment: WSDL, SOAP, RAD, SQL, PL/SQL, JavaScript, DHTML, XHTML, Java Mail, PL/SQL Developer, Toad, POI Reports, Windows XP, Red Hat Linux.
Confidential
Java Developer
Responsibilities:- Responsible for writing functional and technical documents for the modules developed.
- Extensively used J2EE design Patterns.
- Used Agile/Scrum methodology to develop and maintain the project.
- Developed and maintained web services using XMPP and SIP protocols.
- Developed business logic using Spring MVC.
- Developed DAO layer using Hibernate, JPA, and Spring JDBC.
- Used Oracle 10g as the database and used Oracle SQL developer to access the database.
- Used Eclipse Helios for developing the code.
- Used Oracle SQL developer for the writing queries or procedures in SQL.
- Implemented Struts tab libraries for HTML, beans, and tiles for developing User Interfaces.
- Extensively used Soap UI for Unit Testing.
- Involved in Performance Tuning of the application.
- Used Log4J for extensible logging, debugging and error tracing.
- Used Oracle Service Bus for creating the proxy WSDL and then provide that to consumers
- Used JMS with Web Logic Application server.
- Used UNIX scripts for creating a batch processing scheduler for JMS Queue.
- Need to discuss with the client and the project manager regarding the new developments and the errors.
- Documented all the modules and deployed on server in time.
- Involved in Production Support and Maintenance for Application developed in the Red Hat Linux Environment.
Environment: Java 1.5, Spring, Hibernate, XML, XSD, XSLT, WSDL, Web services, XMPP, SIP, JMS, SOAP UI, Eclipse, IBM-UDB, Web logic, Oracle 10g, Oracle SQL developer.
