Hadoop Consultant Resume
San Jose, CA
SUMMARY:
- Over 6+ years of professional IT experience with Big Data Technology including Hadoop/YARN, Pig, Hive, HBase, Cassandra and Spark.
- Hands on experience with Apache Spark, Spark SQL and Spark Streaming.
- Worked with different distributions of Hadoop and Big Data technologies including Hortonworks and Cloudera.
- Expertize in Big Data Hadoop Ecosystem like Flume, Hive, Cassandra, Sqoop, Oozie, ZooKeeper, Kafka etc.
- Well versed with Developing and Implementing MapReduce programs using Java and Python.
- Familiarity with NoSQL databases like HBase and Cassandra.
- Familiarity on real time streaming data with Spark, Kafka.
- Detailed knowledge and experience of Design, Development and Testing Software solutions using Java and J2EE technologies.
- Experience in Database design, Entity relationships, Database analysis, Programming SQL,PL/ SQL, Packages and Triggers in Oracle and SQL Server on Windows and LINUX.
- Strong understanding of Data warehouse concepts, ETL, data modeling experience using Normalization, Business Process Analysis, Reengineering, Dimensional Data modeling, physical & logical data modeling.
- Experience with front - end technologies like HTML, CSS and Javascript.
- Research-oriented, motivated, proactive, self-starter with strong technical, analytical and interpersonal skills.
TECHNICAL SKILLS:
Big Data Technologies: Spark,Hadoop, HDFS, Hive, MapReduce, Pig, Sqoop, Flume, Zookeeper, Cloudera.
Scripting Languages: Python, Shell
Programming Languages: Java, Scala, C, C++
Web Technologies: HTML, J2EE, CSS, JavaScript, Servlets, JSP, XML
Frameworks: Struts, Spring, Hibernate
Application Server: IBM WebSphere Server, Apache Tomcat.
DB Languages: SQL, PL/SQL
Databases /ETL: Oracle 9i/10g/11g
NoSQL Databases: Hbase, Cassandra, ElasticSearch, MongoDB.
Operating Systems: Linux, UNIX
PROFESSIONAL EXPERIENCE:
Confidential, San Jose, CA
Hadoop Consultant
Responsibilities:
- Developed Spark SQL jobs that read data from data lake using Hive, transform and save it in HBase.
- Built Java client that is responsible for receiving XML file using REST call and publishing it to Kafka.
- Built Kafka + Spark Streaming job that is responsible for reading XML file messages from Kafka and transforming it to POJO using JAXB.
- Built Spark + Drools integration that lets us develop Drools rules as part of Spark Streaming job.
- Built HBase DAO’s that responsible for querying data that Drools needs from HBase.
- Built logic to publish output of Drools rules to Kafka for further processing.
Environment: Hadoop, HDFS, Hive, Spark, Spark SQL, Spark Streaming, Kafka, HBase, REST, OpenShift.
Confidential, Walnut Creek, CA
Hadoop Consultant
Responsibilities:
- Developed Sqoop jobs for extracting data from different databases, for both initial and incremental data load
- Developed MapReduce jobs for cleaning up the ingested data, as well as calculating computed fields.
- Designed Hive external tables for storing data extracted using Sqoop.
- Developed Hive jobs for moving data from Avro to ORC format, ORC format was used to speed up the queries
- Created Hive External tables for derived data and loaded the data into tables and query data using HQL for calculating the claim fraud flags.
- Designed Hive External tables with ElasticSearch as Storage format for storing the results of claim flag calculation
- Implemented the workflows using Apache Oozie framework to orchestrate end to end execution.
- Implemented Fair schedulers on the Job tracker to share the resources of the Cluster for the mapreduce jobs given by the users.
- Participated in the Oracle Goldengate POC that would be used for bringing CDC changes to Hadoop using Flume.
- Load log data into HDFS using Flume. Worked extensively in creating MapReduce jobs to power data for search and aggregation.
Environment: Hadoop, MapReduce, HDFS, Flume, Sqoop, Hive, ZooKeeper, Hortonworks, Oozie, ElasticSearch, Sqoop, NoSQL, UNIX/LINUX.
Confidential, Houston, TX
Hadoop Consultant
Responsibilities:
- Obtained the requirement specifications from the SME’s, Business Analysts in the BR, and SR meetings for corporate workplace project. Interacted with the Business users to build the sample report layouts.
- Involved in writing the HLD’s along with the RTM’s tracing back to the corresponding BR’s and SR’s and reviewed them with the Business.
- Load log data into HDFS using Flume.Worked extensively in creating MapReduce jobs to power data for search and aggregation.
- Installed and configured Apache Hadoop and Hive/Pig Ecosystems.
- Created Mapreduce Jobs using Hive/Pig Queries.
- Extensively used Pig for data cleansing.
- Developed the Pig UDFs to pre-process the data for analysis.
- Developed workflow in Oozie to automate the tasks of loading the data into HDFS and pre-processing with Pig and HiveQL.
- Involved in creating Hive tables, loading with data and writing hive queries which will run internally in mapreduce way.
- Cluster coordination services through ZooKeeper.
- Past 5 years TPSS data was collected from Teradata and pushed into HDFS using Sqoop.
Environment: Hadoop, Oracle, Cloudera CDH4, HiveQL, Pig Latin, MapReduce, HDFS, HBase, ZooKeeper, Oozie, Oracle, PL/SQL, Windows, Linux.
Confidential, New York, NY
Hadoop/Big data Engineer
Responsibilities:
- Experience working with Sqoop for importing and exporting data between HDFS and RDBMS systems.
- Designed a data warehouse using Hive. Created partitioned tables in Hive.
- Developed the Hive UDF to pre-process the data for analysis.
- Analyzed the data by performing Hive queries and running Pig scripts to know Artist behavior.
- Worked on data lake concepts, converted all ETL jobs into pig/hive scripts.
- Wrote MapReduce jobs to generate reports for the number of activities created on a particular day, during a dumped from the multiple sources and the output was written back to HDFS.
- Worked on Oozie workflow, cron job.
- Worked with Tableau team in creating Dashboards and built Data visualizations using Tableau and provide analysis on the data.
- Exported analyzed data using Sqoop for generating reports.
- Extensively used Pig for data cleansing. Developed Hive scripts to extract the data from the web server output files.
Environment: Hadoop 1.2, MapReduce, HDFS, Pig, Hive, Sqoop
Confidential, Columbus, OH
Java/J2EE Developer
Responsibilities:
- Involved in requirement gathering, functional and technical specifications.
- Monitoring and fine tuning IDM performance and Enhancements in the self-registration process.
- Developed OMSA GUI using MVC architecture, Core Java, Java Collections, JSP, JDBC, Servlets, ANT and XML within a Windows and UNIX environment.
- Used Java Collection Classes like ArrayList, Vectors, Hash Map and Hash Table.
- Used Design Patterns MVC, Singleton, Factory, Abstract Factory.
- Wrote requirements and detailed design documents, designed architecture for data collection.
- Developed algorithms and coded programs in Java.
- Involved in design and implementation using Core Java, Struts, and JMS
- Performed all types of testing includes Unit testing, Integration and testing environments.
- Worked on a modifying an existing JMS messaging framework for increased loads and performance optimizations.
Environment: JAVA, Design Patterns, Oracle, SQL/ PL SQL,JMS.
