We provide IT Staff Augmentation Services!

Sr. Hadoop Engineer/lead Resume

4.00/5 (Submit Your Rating)

Pleasanton, CA

SUMMARY

  • 15+ years of experience in Analysis, Design and Development of Enterprise, Web and Client Server applications using Java, Bigdata and Business Intelligence technologies.
  • Experienced in development of Bigdata solutions using Cloudera Distribution, Java MapReduce, Apache Spark, Sqoop, Hive, Pig, Apache Solr, ZooKeeper and HBase
  • Expertise in implementation of Concepts like Spark, MapReduce, Hive, Pig, Cloudera Morphline, HBase and Cascading Framework
  • Experienced in planning Hadoop cluster setups and configuring both POC environment and also enterprise environments.
  • Expertise in implementing Data Pipeline Frameworks like Cascading and crunch.
  • Expertise in implementing Design Patterns like Factory Pattern, Singleton, DAO and MVC.
  • Experienced in other programming languages like Python, Scala, C, C++, Shell Scripting
  • Knowledge in using Java profiling tools like VisualVM.
  • Experienced in development of Business Intelligence tools like Tableau, Business Objects and IBM Cognos
  • Involved in using various development IDEs like Eclipse, IntelliJ, Notepad++, TextPad, Sublime Text
  • Expertise in Agile and Waterfall methodologies and a certified Scrum Master.
  • Experienced in unit testing frameworks like Junit, MR unit and Mockito
  • Proficient in Database Systems like Oracle, MySQL, SQL Server and DB2 and a Oracle trained DBA (OCA - Part1 complete)
  • Having Excellent Analytical, Problem Solving, Presentation, communication and interpersonal Skills with ability to interact with any level.
  • Lead a team of developers and can also work as an individual contributor.
  • Expertise in building application roadmaps to help management visualize the tool growth.
  • Expertise in being the consultant to transform user requirements into IT solutions.
  • Training on Data science course on Coursera to help facilitate Data analytics data needs.

PROFESSIONAL EXPERIENCE

Sr. Hadoop Engineer/Lead

Confidential, Pleasanton, CA

Responsibilities:

  • Design the data flows from ingestion to processing and display in Presentation Layer using the Lambda Architecture.
  • Data Ingestion code written using Sqoop, WebCrawler, Flat File ingestion for loading approximately around 80G of data weekly
  • Data ETL Processing code written using Java MapReduce, Apache Spark, Morphlines and Load into Solr
  • UI layer built using Angular JS framework fed from REST Service to read Solr documents
  • Developed Rule Engine API using Java that could be plugged in-memory during the data ETL processing
  • Developed Avro Serialized Schema to Store data in HDFS for efficiency and better development ease.
  • Implemented Hbase-indexer (provided by Cloudera) to help auto indexing of data into Solr and improve cycle time.
  • Lead a team of Senior developers to have the solution developed and deployed
  • Used JENKINS to build and deploy the code in Dev and SIT environments
  • Designed and Implemented REST web services using JAX-RS, Spring REST.
  • Interface point with all the functional teams Release team, infrastructure team, business teams etc.
  • Explore newer tools and technologies to help business gain more productivity. Solr features, QPL support, Cloudera 5.1 to 5.4 upgrade, Java MR to Cascading framework, Hadoop HDFS to HBase etc.
  • As an Additional responsibility, help the SCIF management turn the project into a Scrum style of development. Successfully implemented in Phase 2.0 of the project.

Technology: Java 1.7, Python, CDH4.6.1/5.1/5.3.1/5.4, Sqoop (oraOop), Spark, Hive, Pig, MapReduce, Hbase, Oracle, Solr 4.1/4.4/4.10, Morphline, SQL, Avro, Kerberos, Shell Scripting

Sr. Hadoop/Java Consultant

Confidential, Portland, OR

Responsibilities:

  • Develop/Maintain Avro parser API (Java Spring Framework) to help build the product level dataset.
  • Develop MapReduce jobs to transform the raw sports data into product dataset using AvroInputFormat and AvroSerde in Hive tables. Approximately 20TB of data for 2013 ingested for processing.
  • Develop FLUME stream for twitter data for a POC purpose to ingest Twitter data.
  • Work with the Analysts to build the datasets for testing and training their Models .
  • Work with all the dependent teams (source teams, project management, Hadoop Admin) to resolve issues for the users.
  • Develop UDFs in Hive for custom function development (geodistance, deviceType, etc. ).
  • Development of Shell Scripts to build a workflow as mandated by the organization.
  • Development of SQOOP jobs to ingest approx. 200GB initial and 20GB per day.
  • Development of Data pipeline using Cascading Framework to replace existing MapReduce Jobs.

Technology: Java 1.7, CDH4.5, Sqoop, Hive, MapReduce v1, Avro serialization,, Oracle, Flume. Cascading Framework, Shell Scripts, Pig

We'd love your feedback!