Hadoop/spark Developer Resume
Santa Maria, CA
PROFESSIONAL SUMMARY
- 8 years of professional IT experience in analysis, design, architecture, development, testing and implementation of Hadoop, Bigdata Technologies, Data Warehousing, and AWSon Object Oriented Programming.
- Expertise in Hadoop, MapReduce, YARN, Hive, Pig, Sqoop, Kafka, Spark - Scala, and Spark Streaming.
- Hands on experience with AWS components like EC2, S3, Data Pipeline, RDS, RedShift and EMR.
- In depth knowledge on Hadoop Architecture and various components such as HDFS, Job Tracker, Name node, Data Node and MapReduce concepts.
- Good understanding of HDFS Designs, Daemons, HDFS High Availability (HA).
- Good understanding and working experience on Hadoop Distributions like Cloudera.
- Worked on NoSQL databases like MongoDB,HBase,Cassandra.
- Experience in analyzing data using HiveQL,Pig Latin and custom MapReduce programs in Java.
- Having Good Working Expertise on handling Multi Terabytes of structured and unstructured data on significantly big Cluster Environment
- Hands on experience in working with Flume to load the log data from multiple sources directly into HDFS.
- Experience in analyzing large scale data to identify new analytics, insights, trends, and relationships with a strong focus on data clustering.
- Have good knowledge on Data Center OS Mesos, and Docker.
- Hands on experience using business intelligence tools like Tableau, Cognos and Qlik View.
- Proficient in developing web based applications and client server distributed architecture applications in Java/J2EE technologies using Object Oriented Methodology.
- Strong Knowledge on full Software Development life cycle -Software analysis, design, architecture, development, and maintenance.
- Worked on relative ease with different working strategies like Agile, Waterfall, Scrum, and Test-Driven Development(TDD) methodologies.
- Excellent experience in designing and developing Enterprise Applications for J2EE platform using Servlets, JSP, Struts, Spring, Hibernate and Web services.
- Expertise in developing web services with XML based protocols such as SOAP and WSDL.
- Learning new technologies will allow for more effective design and implementation.
- Intellectual capacity to grasp new situations quickly and problem-solving skills.
- Excellent communication, interpersonal skills, proactive and a good team player along with a can-do attitude and good written skills.
TECHNICAL SKILLS
Big Data Technologies: HDFS, Hive, Map Reduce, Pig, Sqoop, Oozie, Zookeeper, YARN, Spark.
Scripting Languages: Shell, Python, Scala.
Programming Languages: CoreJava, C++, C, SQL.
Front End Technologies: HTML, XHTML, CSS, XML, JavaScript, AJAX, Servlets, JSP.
Web Services: AWS EC2, S3, Data Pipeline, RDS, RedShift, EMR, Dynamo DB, SOAP, and Rest.
Application Servers: Apache Tomcat, WebLogic Server, WebSphere, JBoss.
Databases: Oracle 11g, MySQL, IBM DB2
DW & BI Tools: DataStage, Tableau, Cognos, Qlik View, and Qlik Sense.
NoSQL Databases: HBase, MongoDB, and Cassandra.
IDE: Eclipse, NetBeans, JBuilder.
Operating Systems: Linux, UNIX, MAC, Windows NT / 98 /2000/ XP / Vista, Windows 7, Windows 8.
PROFESSIONAL EXPERIENCE
Confidential, Santa Maria, CA
Hadoop/Spark Developer
Responsibilities:
- Prepared an ETL framework with the help of sqoop, pig and hive to be able to frequently bring in data from the source and make it available for consumption.
- Worked on importing and exporting data from Oracle and DB2 into HDFSusing Sqoop.
- Developed data pipeline using Flume to ingest customer behavioral data and financial histories into HDFS for analysis.
- Scheduled a workflow to import the weekly transactions in the revenue department from RDBMS database using Oozie.
- Built wrapper shell scripts to hold theseOozieworkflow.
- Developed PIG Latin scripts to transform the log data files and load processed datasets into HDFS.
- Used Pig as ETL tool to do transformations, event joins and some pre-aggregations before storing the data onto HDFS.
- Hands on experience with NoSQL databases like Cassandra for POC (proof of concept) in storing URL's and images.
- Developedhive UDF for functions that were not preexisting in Hive like the rank etc.
- Created ExternalHive tables and involved in data loading and writing Hive UDFs.
- Migrated iterative map reduce programs into Spark transformations using Scala.
- Experienced in performance tuning of Spark Applications for setting right Batch Interval time, correct level of Parallelism and memory tuning.
- Optimizing of existing word2vec algorithms in Hadoop using Spark Context, Spark-SQL, Data Frames, and Pair RDD's in development of Chatbot using OpenNLP and Word2Vec.
- Experienced in handling large datasets using Partitions, Spark in Memory capabilities, Broadcasts in Spark, Effective & efficient Joins, Transformations and other during ingestion process itself.
- Created concurrent access for hive tables with shared and exclusive locking that can be enabled in hive with the help of Zookeeper implementation in the cluster.
- Build Tez source code and configured on Hive and achieved very good responsive time (<1 min) while running the huge Hive queries which used to take longer time (> 30 mins)
- Continuous monitoring and managing the Hadoop cluster through Cloudera Manager.
- Wrote the shell scripts to monitor the health check of Hadoop daemon services and respond accordingly to any warning or failure conditions.
- Developed Unit test cases using MRunit for map reduce code.
Environment:Hadoop, HDFS, MapReduce, Hive, Pig, HBase, Sqoop, Spark, Oozie, Zookeeper, RDBMS/DB, MySQL, CSV, AVRO data files.
Confidential, Atlanta, GA
Hadoop Developer
Responsibilities:
- Responsible for loading the customer's data and event logs from Oracle database, Teradata into HDFS using Sqoop
- Involved in initiating and successfully completing Proof of Concept on SQOOP for Pre-Processing, Increased Reliability and Ease of Scalability over traditional Oracle database.
- End-to-end performance tuning of Hadoop clusters and Hadoop MapReduce routines against very large data sets.
- Developed the Pig UDF'S to pre-process the data for analysis.
- Involved in loading data from LINUX file system to HDFS.
- Importing and exporting data into HDFS and Hive using Sqoop and Flume.
- Proficient in using Cloudera Manager, an end to end tool to manage Hadoop operations.
- Developed MapReduce jobs for Log Analysis, Recommendation and Analytics.
- Wrote MapReduce jobs to generate reports for the number of activities created on a particular day, during a dumped from the multiple sources and the output was written back to HDFS
- Reviewed the HDFS usage and system design for future scalability and fault-tolerance.
- Installed and configured Hadoop HDFS, MapReduce, Pig, Hive, and Sqoop.
- Wrote Pig Scripts to generate MapReduce jobs and performed ETL procedures on the data in HDFS.
- Exported analyzed data to HDFS using Sqoop for generating reports.
- Used MapReduce and Sqoop to load, aggregate, store and analyze web log data from different web servers.
- Developed Hive queries for the analysts.
- Cluster co-ordination services through Zookeeper.
- Written the Spouts and Bolts after collecting the real stream customer data from Kafka broker to process and store into HBASE.
- Analyze the log files and process through Flume
- Experience in optimization of MapReduce algorithm using combiners and partitions to deliver the best results and worked on Application performance optimization.
Environment: CDH4, MapReduce, HDFS, Hive, Pig, Sqoop, Linux, XML, MySQL, MySQL Workbench,PL/SQL, SQL connector
Confidential, Chicago, IL
Hadoop/ETL Developer
Responsibilities:
- Responsible for managing data from multiple sources.
- Involved in the migration part of the project for 30 sources.
- Worked with IBM data extraction application Data stage for ETL purpose to get the data on Edge Node.
- Developed CRON jobs to write the input data files to HDFS location and Archive location.
- Developed Map Reduce applications for the schema validation and Row Count Validation.
- Assisted in exporting analyzed data to relational databases using Sqoop.
- Experienced in importing and exporting data into HDFS and assisted in exporting analyzed data to RDBMS using SQOOP.
- Once Schema and Row Count Validation is done, written MR jobs to create Avro Schema.
- Developed Applications to convert .dat files to Avro data format.
- Written MR jobs to create super set schema from different Avro schemas.
- Hive tables have been created from the super set schema.
- To run multiple MapReduce and Hive jobs installed and used Oozie Workflow engine.
- Participated in white board sessions to get the task requirements.
- Handled importing of data from various data sources, performed transformations using Hive, MapReduce, and loaded data into HDFS.
- Extracted the data from Teradata into HDFS using Sqoop.
- Analyzed the data by performing Hive queries and running Pig scripts to know user behavior like shopping enthusiasts, travelers, music lovers etc.
- Wrote REST Web services to expose the business methods to external services.
- Exported the patterns analyzed back into Teradata using Sqoop.
- Continuous monitoring and managing the Hadoop cluster through Cloudera Manager.
- Installed Oozie workflow engine to run multiple Hive.
Environment: Hadoop, MapReduce, HDFS, Hive, Flume, Sqoop, Cloudera, Oozie, Data Stage, Java Cron Jobs, UNIX Scripts.
Confidential
Java/J2EEDeveloper
Responsibilities:
- Involved in Analysis, Designing, Development and Testing phases of the application.
- Was involved in creation and maintenance of the backend services using Spring, Hibernate, SQLServer and Oracle.
- Developed Web pages using JSPs with Tag libraries, HTML, and JavaScript.
- Writing J2EE code using Spring, hibernate to upload input CSV files for credit risk data.
- Implemented Dependency Injection (IOC) feature of spring framework to inject dependency into objects and AOP is used for Logging.
- Designed and developed persistence layer build on ORM framework and developed it using Hibernate
- Implemented various Design patterns like Business Delegate, Data Transfer Objects DTO, Service locator, Session Facade and Data Access Objects DAO patterns.
- Involved in writing SQL, Stored procedure, and PL/SQL for back end. Used Views and Functions at the Oracle Database end.
- Developed various documents within the application using XML by using Eclipse as IDE tool.
- Developed SOAP requests to interact with billing schedule system.
- Used Web Services (SOAP & WSDL) to exchange data between Server part and Comercia Bank.
- Integrating and deploying the application on WebLogic application server using ANT.
- Developed user interfaces for presenting the expense reports, transaction details using JSP, XML, HTML, and Java Script.
- Used Log4J for logging the application exceptions and debugging statements.
- Proficient in doing Object Oriented Design using UML-Rational Rose
Environment: Java, JSP, Servlets, Web Sphere Application Server, Eclipse, Java Script, Web Services (SOAP & WSDL), Microsoft VSS, Oracle, PL/SQL and JDBC.
Confidential
Java/J2EEDeveloper
Responsibilities:
- Utilized the base UML methodologies and Use cases modeled by architects to develop the front-end interface. The class, sequence and state diagrams were developed using Rational Rose and Microsoft Visio.
- Designed application using MVC design pattern.
- Developed front-end user interface modules by using HTML, XML, Java AWT, and Swing.
- Front-end validations of user requests carried out using Java Script.
- Designed and developed the interacting JSPs and Servlets for modules like User Authentication and Summary Display.
- Designed and developed Entity/Session EJB components for the primary modules.
- Java Mail was used to notify the user of the status and completion of the request.
- Developed Stored Procedures on Oracle 8i.
- Implemented Queries using SQL (database triggers and functions).
- JDBC was used to interface the web-tier components on the J2EE server with the relational database.
- Performed client-side validations using JavaScript and AJAX.
- Coded the JUnit test cases and suite for the system.
- Project coordination with other Development teams, System managers and web master and developed good working environment
Environment: Java 1.5, Servlets, J2EE 1.4, JDBC, Oracle, 10g, PL SQL, HTML, JSP, Eclipse, UNIX
