We provide IT Staff Augmentation Services!

Sr. Big Data/hadoop Developer Resume

0/5 (Submit Your Rating)

SUMMARY

  • Over 8 years of Software Development experience, 3 years of Big Data experience in ingestion, storage, querying, processing and analysis
  • Expertise in Hadoop ecosystem - HDFS, YARN, Pig, HBase, Spark and Hive for data analysis, Sqoop for data migration, Flume for data ingestion, Oozie for scheduling and Zookeeper for coordinating cluster resources.
  • Experience on Amazon cloud components AWS - EC2, EMR, S3.
  • Experience in providing design architecture for Big Data solutions.
  • Strong Data Warehouse ETL experience in Advertising, Banking & Insurance financial domain.
  • Involved in designing Hive schemas, using performance tuning techniques like partitioning, bucketing.
  • Optimized HiveQL/ pig scripts by using execution engine like Tez, Spark.
  • Migrating EDW (Enterprise Data Warehouse) into Big Data and implemented Star Schema in Big Data.
  • Expertise in implementing HBase schemas with optimized Row- key design to avoid Hot-spotting.
  • Good understanding on Spark SQL, Spark Transformation Engine and Spark Streaming.
  • Experience working with Scala.
  • Loaded streaming log data from various web servers into HDFS using Flume.
  • Exposed HBase tables to web applications with REST web services.
  • Used Apache Solr to create full text searches for more than 9 million rows.
  • Experience in integrating Hive server with visualization tools like Tableau, Qlikview, Informatica using ODBC driver.
  • Experience in building Pig scripts to extract, transform and load different file formats- JSON, TXT, XML data onto HDFS, HBase, and Hive for data processing.
  • Used SFTP to transfer the files to server.
  • Analyzed the data using Hive queries and running Pig scripts to study customer behavior.
  • Developed Pig UDF'S to pre-process the data for analysis.
  • Good understanding of Java Object Oriented Concepts and development of multi-tier enterprise web applications.

TECHNICAL SKILLS

Hadoop Ecosystem: HDFS, YARN, Spark, Pig, Hive, HBase, Scala, Oozie, Flume, Solr, Hue, AmbariDistribution/Cluster Amazon AWS, IBM BigInsights 4.1, Hortonworks, Cloudera CDH4

Database: Vertica, MySQL, Oracle

Languages: Java, C, C++, JSP, Shell Script

Web Technologies: HTML5, CSS3, JavaScript, jQuery, XML, XHTML.

Servers: Putty, WebSphere, WebLogic, JBoss, Apache Tomcat.

Operating Systems: Macintosh, Linux, Windows

PROFESSIONAL EXPERIENCE

Confidential

Sr. Big Data/Hadoop Developer

Responsibilities:

  • Prepared an ETL framework with the help of sqoop, pig and hive to be able to frequently bring in data from the source and make it available for consumption.
  • Loaded the data from Vertica database to HDFS and Amazon S3.
  • Responsible for creating scripts/jobs to migrate data from Amazon S3 to Hadoop platform and vice versa.
  • Spark streaming collects the data from Kafka in near real time and performs necessary transformations and aggregations on the fly to build the common learner data model.
  • Wrote Pig Scripts to generate MapReduce jobs and performed ETL procedures on the data in HDFS.
  • Created 20 buckets for each Hive table based on clustering by client Id for better performance (optimization) while updating the tables.
  • Experience in streaming the data between Kafka and other databases like RDBMS and NoSQL.
  • Used Spark with YARN and got performance results compared with MapReduce.
  • Wrote shell scripts to run the Cron jobs to automate the data migration process from external servers and FTP sites.
  • Troubleshooting and analyzing Hadoop clusters for job failures.
  • Involved in gathering the requirements, designing, development and testing.
  • Experience in deploying the application in production and Disaster Recovery (DR) servers and also analyzing them in cases of job failures.
  • Collaborated with the Internal/Client BAs in understanding the requirement and architect a data flow system.
  • Configured High Availability on the cluster.
  • Supported code/design analysis, strategy development and project planning.
  • Analyzed the data using Map Reduce, Pig, Hive and produce summary results from Hadoop to downstream systems.
  • Loaded the data into Hive partitioned tables, on the basis of AOL vendors.
  • Load and transform large sets of structured, semi structured and unstructured data using Big Data concepts.
  • Optimized hive scripts to use HDFS efficiently by using various compression mechanisms.
  • Writing UDF/MapReduce jobs depending on the specific requirement.
  • Manage and review Hadooplog files, File system management and monitoring.
  • Tested raw data and executed performance scripts.
  • Documented the systems processes and procedures for future references.
Environment: ETL, Vertica, Oracle, HDFS, YARN, Amazon AWS, Pig, Hive, NoSQL- HBase, Sqoop, Spark, Scala, Khafka, Sqoop, Shell Scripts, Cron, Job Execution Framework (JEF).

Confidential

Sr. Hadoop Developer

Responsibilities:

  • Configured Hadoop components including Hive, Pig, HBase, Spark, Sqoop, Oozie and Hue in the client environment.
  • Stored Solr indexes in HDFS.
  • Index documents in HDFS using Solr Hadoop connectors.
  • Responsible to manage data coming from different sources and involved in HDFS maintenance and loading of structured and unstructured data.
  • Developed MapReduce programs to parse the raw data, populate staging tables and store the refined data in partitioned tables in HDFS.
  • Worked on the backend using Scala and Spark to perform several aggregation logics.
  • Created Hive queries that helped market analysts spot emerging trends by comparing fresh data with HDFS reference tables and historical metrics.
  • Enabled speedy reviews and first mover advantages by defining the job flow in Oozie to automate data loading into the Hadoop Distributed File System and PIG to pre-process the data.
  • Designed HBase schema to avoid Hotspotting and exposed the data from HBase tables to REST API on UI.
  • DevelopedPig scripts to transform raw datafrom several data sources into forming baseline data and loaded the data into HBase tables.
  • Involved in creating POCs to ingest and process streaming data using Spark and HDFS.
  • Used Flume to collect, aggregate, and store the log data from different web servers.
  • Developed Shell Scripts to automate the batch processing and processed the daily jobs through Maestro scheduler.
  • Provided design recommendations and thought leadership to sponsors/stakeholders that improved review processes and resolved technical problems.
  • Co-ordinate with the offshore team and cross-functional teams to ensure that applications are properly tested, configured, and deployed.
  • Used Tableau for visualizing and to generate reports.

Environment: HDFS, YARN, Pig, Hive, HBase, Spark, Scala, Solr, Sqoop, Flume, Oozie, Shell Scripts.

Confidential

Hadoop Developer

Responsibilities:

  • Installation and configuration of Hadoop/HDFS multi-node cluster.
  • Involved in installing, configuring and managing Hadoop Ecosystem components like Hive, Pig, Sqoop.
  • Developed Map Reduce programs, PIG, Hive scripts to clean and filter data on the cluster and store them in HDFS.
  • Executing/Monitoring MR jobs on cluster.
  • Exported the business required information to RDBMS using Sqoop to make the data available for BI team to generate reports based on data.
  • Responsible for creating Hive tables, loading data and writing hive queries.
  • Designed HBase tables for faster results.
  • Developing Pig/Hive scripts depending on the business logic complexity.
  • Created partitioned tables in Hive.
  • Executed Oozie workflows to run multiple Hive and Pig jobs.
  • Involved in loading data from UNIX file system to HDFS.
  • Writing Custom writable classes for Hadoopserialization and De serialization.

Environment: Hadoop, Map Reduce, Hive, Pig, HBase, Sqoop, Oozie.

Confidential

Hadoop Developer

Responsibilities:

  • Migrating the needed data from MySQL in to HDFS using Sqoop and importing various formats of flat files into HDFS.
  • Mainly worked on Hive queries to categorize data of different claims
  • Integrated the hive warehouse with HBase for information sharing among teams.
  • Written customized Hive UDFs in Java where the functionality is too complex.
  • Designed and created Hive external tables using shared meta-store and supported partitioning, dynamic partitioning for faster data retrieval.
  • Developed the Sqoop scripts in order to make the interaction between Pig and MySQL Database.
  • EJB session Beans being used to interact with Database using the JPA.
  • HiveQL scripts to create, load, and query tables for extracting the summarized information
  • Implemented modules using Core Java APIs, Java collection, Threads and integrating the module.
  • Used Avro's the file storage format to save disk storage space.
  • Maintained System integrity of sub components primarily HDFS, MR, HBase and Hive.
  • Monitored System health and logs and respond accordingly to any warning or failure conditions.

Environment: Apache Hadoop, Shell scripting, HDFS, Hive, Map Reduce, HBase, Java, Pig, Sqoop, Cloudera CDH4, MySQL, Tableau, Avro, Spring, EJB, XML, Java Collections, REST

Confidential

Java Developer

Responsibilities:

  • Involved in the complete SDLC software development life cycle of the application from requirement analysis to testing.
  • Developed the modules based on struts MVC Architecture.
  • Worked with various types of controllers like simple form controller, Abstract Controller and Controller Interface etc.
  • Developed UI modules using HTML, JSP, JavaScript and CSS.
  • Involved in writing and executing queries in MySQL.
  • Build test cases and performed unit testing.
  • Implemented code for validating the input fields and displaying the error messages.
  • Provided Technical support for production environments resolving the issues, analyzing the defects, providing and implementing the solution defects.
  • Developed coded, tested, debugged and deployed JSPs and Servlets for the input and output forms on the web browsers.
  • Database Modification using SQL, PL/SQL, Stored procedures, triggers, Views in Oracle9i.
  • Experience in going through bug queue, analyzing and fixing bugs, escalation of bugs.

Environment: Java, MVC, HTML, CSS, JavaScript, JSP, MySQL, Oracle, JDBC, Tomcat Web Server.

Confidential

Java Developer

Responsibilities:

  • Worked on Full Cycle of Software Development from Analysis through Design, Development, Testing, Integration, Deployment.
  • Extensively used Spring MVC Framework.
  • Designed and developed User Interface of application modules using HTML, JSP, CSS, JavaScript, jQuery and AJAX.
  • Design and development of modules using MVC.
  • Developed Struts Action Classes, Action Forms and performed Action mapping using Struts framework.
  • Worked on XML, XSLT, XPATH, DOM, SAX.
  • Created XML-SOAP Web Services to provide partner systems required information.
  • Used Websphere Application Server for deploying the application.
  • Used Rational Application Developer (RAD) for developing the application.
  • Involved in storing paper forms into IBM Content Manager after converting them to images (.jpeg, .tiff, etc) and also storing Electronic forms.
  • Prepared Unit Test Plan & performed Unit Testing using JUnit.
  • Actively involved in Walkthroughs and Peer Reviews.
  • Used JIRA for bug tracking.
  • Used SVN as version control system for the source code.

Environment: Java, JEE, Oracle, JIRA, Servlets, JSP, Struts, Spring, XML, UML, JBOSS, RAD, SVN

We'd love your feedback!