Sr. Big Data/hadoop Developer Resume
SUMMARY
- Over 8 years of Software Development experience, 3 years of Big Data experience in ingestion, storage, querying, processing and analysis
- Expertise in Hadoop ecosystem - HDFS, YARN, Pig, HBase, Spark and Hive for data analysis, Sqoop for data migration, Flume for data ingestion, Oozie for scheduling and Zookeeper for coordinating cluster resources.
- Experience on Amazon cloud components AWS - EC2, EMR, S3.
- Experience in providing design architecture for Big Data solutions.
- Strong Data Warehouse ETL experience in Advertising, Banking & Insurance financial domain.
- Involved in designing Hive schemas, using performance tuning techniques like partitioning, bucketing.
- Optimized HiveQL/ pig scripts by using execution engine like Tez, Spark.
- Migrating EDW (Enterprise Data Warehouse) into Big Data and implemented Star Schema in Big Data.
- Expertise in implementing HBase schemas with optimized Row- key design to avoid Hot-spotting.
- Good understanding on Spark SQL, Spark Transformation Engine and Spark Streaming.
- Experience working with Scala.
- Loaded streaming log data from various web servers into HDFS using Flume.
- Exposed HBase tables to web applications with REST web services.
- Used Apache Solr to create full text searches for more than 9 million rows.
- Experience in integrating Hive server with visualization tools like Tableau, Qlikview, Informatica using ODBC driver.
- Experience in building Pig scripts to extract, transform and load different file formats- JSON, TXT, XML data onto HDFS, HBase, and Hive for data processing.
- Used SFTP to transfer the files to server.
- Analyzed the data using Hive queries and running Pig scripts to study customer behavior.
- Developed Pig UDF'S to pre-process the data for analysis.
- Good understanding of Java Object Oriented Concepts and development of multi-tier enterprise web applications.
TECHNICAL SKILLS
Hadoop Ecosystem: HDFS, YARN, Spark, Pig, Hive, HBase, Scala, Oozie, Flume, Solr, Hue, AmbariDistribution/Cluster Amazon AWS, IBM BigInsights 4.1, Hortonworks, Cloudera CDH4
Database: Vertica, MySQL, Oracle
Languages: Java, C, C++, JSP, Shell Script
Web Technologies: HTML5, CSS3, JavaScript, jQuery, XML, XHTML.
Servers: Putty, WebSphere, WebLogic, JBoss, Apache Tomcat.
Operating Systems: Macintosh, Linux, Windows
PROFESSIONAL EXPERIENCE
Confidential
Sr. Big Data/Hadoop Developer
Responsibilities:
- Prepared an ETL framework with the help of sqoop, pig and hive to be able to frequently bring in data from the source and make it available for consumption.
- Loaded the data from Vertica database to HDFS and Amazon S3.
- Responsible for creating scripts/jobs to migrate data from Amazon S3 to Hadoop platform and vice versa.
- Spark streaming collects the data from Kafka in near real time and performs necessary transformations and aggregations on the fly to build the common learner data model.
- Wrote Pig Scripts to generate MapReduce jobs and performed ETL procedures on the data in HDFS.
- Created 20 buckets for each Hive table based on clustering by client Id for better performance (optimization) while updating the tables.
- Experience in streaming the data between Kafka and other databases like RDBMS and NoSQL.
- Used Spark with YARN and got performance results compared with MapReduce.
- Wrote shell scripts to run the Cron jobs to automate the data migration process from external servers and FTP sites.
- Troubleshooting and analyzing Hadoop clusters for job failures.
- Involved in gathering the requirements, designing, development and testing.
- Experience in deploying the application in production and Disaster Recovery (DR) servers and also analyzing them in cases of job failures.
- Collaborated with the Internal/Client BAs in understanding the requirement and architect a data flow system.
- Configured High Availability on the cluster.
- Supported code/design analysis, strategy development and project planning.
- Analyzed the data using Map Reduce, Pig, Hive and produce summary results from Hadoop to downstream systems.
- Loaded the data into Hive partitioned tables, on the basis of AOL vendors.
- Load and transform large sets of structured, semi structured and unstructured data using Big Data concepts.
- Optimized hive scripts to use HDFS efficiently by using various compression mechanisms.
- Writing UDF/MapReduce jobs depending on the specific requirement.
- Manage and review Hadooplog files, File system management and monitoring.
- Tested raw data and executed performance scripts.
- Documented the systems processes and procedures for future references.
Confidential
Sr. Hadoop Developer
Responsibilities:
- Configured Hadoop components including Hive, Pig, HBase, Spark, Sqoop, Oozie and Hue in the client environment.
- Stored Solr indexes in HDFS.
- Index documents in HDFS using Solr Hadoop connectors.
- Responsible to manage data coming from different sources and involved in HDFS maintenance and loading of structured and unstructured data.
- Developed MapReduce programs to parse the raw data, populate staging tables and store the refined data in partitioned tables in HDFS.
- Worked on the backend using Scala and Spark to perform several aggregation logics.
- Created Hive queries that helped market analysts spot emerging trends by comparing fresh data with HDFS reference tables and historical metrics.
- Enabled speedy reviews and first mover advantages by defining the job flow in Oozie to automate data loading into the Hadoop Distributed File System and PIG to pre-process the data.
- Designed HBase schema to avoid Hotspotting and exposed the data from HBase tables to REST API on UI.
- DevelopedPig scripts to transform raw datafrom several data sources into forming baseline data and loaded the data into HBase tables.
- Involved in creating POCs to ingest and process streaming data using Spark and HDFS.
- Used Flume to collect, aggregate, and store the log data from different web servers.
- Developed Shell Scripts to automate the batch processing and processed the daily jobs through Maestro scheduler.
- Provided design recommendations and thought leadership to sponsors/stakeholders that improved review processes and resolved technical problems.
- Co-ordinate with the offshore team and cross-functional teams to ensure that applications are properly tested, configured, and deployed.
- Used Tableau for visualizing and to generate reports.
Environment: HDFS, YARN, Pig, Hive, HBase, Spark, Scala, Solr, Sqoop, Flume, Oozie, Shell Scripts.
Confidential
Hadoop Developer
Responsibilities:
- Installation and configuration of Hadoop/HDFS multi-node cluster.
- Involved in installing, configuring and managing Hadoop Ecosystem components like Hive, Pig, Sqoop.
- Developed Map Reduce programs, PIG, Hive scripts to clean and filter data on the cluster and store them in HDFS.
- Executing/Monitoring MR jobs on cluster.
- Exported the business required information to RDBMS using Sqoop to make the data available for BI team to generate reports based on data.
- Responsible for creating Hive tables, loading data and writing hive queries.
- Designed HBase tables for faster results.
- Developing Pig/Hive scripts depending on the business logic complexity.
- Created partitioned tables in Hive.
- Executed Oozie workflows to run multiple Hive and Pig jobs.
- Involved in loading data from UNIX file system to HDFS.
- Writing Custom writable classes for Hadoopserialization and De serialization.
Environment: Hadoop, Map Reduce, Hive, Pig, HBase, Sqoop, Oozie.
Confidential
Hadoop Developer
Responsibilities:
- Migrating the needed data from MySQL in to HDFS using Sqoop and importing various formats of flat files into HDFS.
- Mainly worked on Hive queries to categorize data of different claims
- Integrated the hive warehouse with HBase for information sharing among teams.
- Written customized Hive UDFs in Java where the functionality is too complex.
- Designed and created Hive external tables using shared meta-store and supported partitioning, dynamic partitioning for faster data retrieval.
- Developed the Sqoop scripts in order to make the interaction between Pig and MySQL Database.
- EJB session Beans being used to interact with Database using the JPA.
- HiveQL scripts to create, load, and query tables for extracting the summarized information
- Implemented modules using Core Java APIs, Java collection, Threads and integrating the module.
- Used Avro's the file storage format to save disk storage space.
- Maintained System integrity of sub components primarily HDFS, MR, HBase and Hive.
- Monitored System health and logs and respond accordingly to any warning or failure conditions.
Environment: Apache Hadoop, Shell scripting, HDFS, Hive, Map Reduce, HBase, Java, Pig, Sqoop, Cloudera CDH4, MySQL, Tableau, Avro, Spring, EJB, XML, Java Collections, REST
Confidential
Java Developer
Responsibilities:
- Involved in the complete SDLC software development life cycle of the application from requirement analysis to testing.
- Developed the modules based on struts MVC Architecture.
- Worked with various types of controllers like simple form controller, Abstract Controller and Controller Interface etc.
- Developed UI modules using HTML, JSP, JavaScript and CSS.
- Involved in writing and executing queries in MySQL.
- Build test cases and performed unit testing.
- Implemented code for validating the input fields and displaying the error messages.
- Provided Technical support for production environments resolving the issues, analyzing the defects, providing and implementing the solution defects.
- Developed coded, tested, debugged and deployed JSPs and Servlets for the input and output forms on the web browsers.
- Database Modification using SQL, PL/SQL, Stored procedures, triggers, Views in Oracle9i.
- Experience in going through bug queue, analyzing and fixing bugs, escalation of bugs.
Environment: Java, MVC, HTML, CSS, JavaScript, JSP, MySQL, Oracle, JDBC, Tomcat Web Server.
Confidential
Java Developer
Responsibilities:
- Worked on Full Cycle of Software Development from Analysis through Design, Development, Testing, Integration, Deployment.
- Extensively used Spring MVC Framework.
- Designed and developed User Interface of application modules using HTML, JSP, CSS, JavaScript, jQuery and AJAX.
- Design and development of modules using MVC.
- Developed Struts Action Classes, Action Forms and performed Action mapping using Struts framework.
- Worked on XML, XSLT, XPATH, DOM, SAX.
- Created XML-SOAP Web Services to provide partner systems required information.
- Used Websphere Application Server for deploying the application.
- Used Rational Application Developer (RAD) for developing the application.
- Involved in storing paper forms into IBM Content Manager after converting them to images (.jpeg, .tiff, etc) and also storing Electronic forms.
- Prepared Unit Test Plan & performed Unit Testing using JUnit.
- Actively involved in Walkthroughs and Peer Reviews.
- Used JIRA for bug tracking.
- Used SVN as version control system for the source code.
Environment: Java, JEE, Oracle, JIRA, Servlets, JSP, Struts, Spring, XML, UML, JBOSS, RAD, SVN
