We provide IT Staff Augmentation Services!

Sr. Hadoop/bigdata Consultant Resume

4.00/5 (Submit Your Rating)

Sacramento, CA

SUMMARY

  • 8+ years of IT experience along with 4+ years of Hadoop experience.
  • Experienced Big Data Engineer with different Hadoop Distributed File Systems and Eco System MapReduce, Pig, Hive, Spark, Scala, Sqoop, Oozie
  • Good understanding of Hadoop Distributed File System and MapReduce.
  • Experienced in installing, configuring Hadoop multi node cluster in Amazon EC2.
  • Work experience in different layers of Hadoop Framework - Storage layer (HDFS), Analysis Layer (Pig and Hive), Engineering Layer (Jobs and Workflows)
  • Expertise in deployment of Hadoop, Yarn,Sparkand Storm integration with Cassandra and Kafka etc.
  • Experience in analyzing data using Pig Latin and Hive QL.
  • Knowledge in job/workflow scheduling and monitoring tools like Oozie & Zookeeper.
  • Importing & exporting data from existing databases that provide SQL interfaces using ETL tool Sqoop.
  • Strong in developing, executing and scheduling Pig UDF’s for data processing
  • Experienced in working with different data sources like Flat files, Spreadsheet files, log files and Databases.
  • Experience on Spark and Scala.
  • Developed analytical components using Scala, Spark.
  • Implemented POC to migrate map reduce jobs into Spark RDD transformations using Scala.
  • Developed Spark applications using Scala for easyHadooptransitions.
  • Developed custom aggregate functions using SparkSQL and performed interactive querying.
  • Implemented Spark using Scala and SparkSQL for faster testing and processing of data.
  • Experience in monitoring and managing Hadoop cluster using Hortonworks
  • Expertise with managing and reviewing Hadoop log files.
  • Background with traditional databases such as Oracle, SQL Server, MySQL.
  • Expertise with NoSQL databases such as HBase.
  • Experience in Elastic search technologies in creating custom Lucene/Solr Query components.
  • Well experienced and possess strong knowledge in Unix Shell Scripting
  • Ability to work independently and in a group with effective communication and quantitative skills.
  • Expertise in implementing Spark and Scala application using higher order functions for both batch and interactive analysis requirement
  • Good working experience in using Spark SQL to manipulate Data Frames.
  • Experiencein working with Spark tools like RDD transformations, spark SQL.
  • Proficiency in formulating a strategic vision and a tactical roadmap to address client's critical Business
  • Intelligence/ Analytics needs in conformance with overall corporate objectives.
  • Technical evangelist skilled at developing new applications on Hadoop according to business needs, andconvert existing applications to Hadoop environment.

TECHNICAL SKILLS

Hadoop/Big Data: CDH4.4, HDFS, MapReduce2, Hive, Pig, HBase Zookeeper, Sqoop, Oozie, Flume, Storm, Spark and Scala

No-SQL Databases: Hbase

Programming Languages: Scala, Java, Pig Latin, HiveQL, Unix shell scripts, SQL, PLSQL.

Operating System: Windows, Unix, Linux, AIX 5.3/7.1

Relational Database: Oracle … DB2, SQL Server 2008, SQL Server 2012, MySQL

Tools: and IDE: Eclipse, NetBeans, IntelliJ

PROFESSIONAL EXPERIENCE

Confidential, Sacramento, CA

Sr. Hadoop/Bigdata Consultant

Responsibilities:

  • Writing Scala User-Defined Functions (UDFs) to solve the business requirements.
  • Creating the Case Classes.
  • Working with the Data Frames and RDD’s.
  • Parsing the JSONObjects into flatten formats using Scala.
  • Creating the tables in Hive and integrating data between Hive &Spark.
  • Worked on the core and Spark SQL modules of Spark extensively using programming languages likeScala.
  • Hands on experience with Spark Scala programming and good understanding of its 'In Memory' processing capability.
  • Worked on creating the RDD's, DF's for the required input data and performed the data transformations using Spark Scala.
  • Experience in Kafka Producer, Consumer, Brokers.
  • Experienced with batch processing of data sources using Apache Spark.
  • Experienced in working with RDDs.
  • Experience in Writing the Scala functions, procedures, Constructors and Traits.
  • Maintenance and Monitoring of Cassandra cluster using Opscenter and Node tool.
  • Cassandra data modeling and Design.
  • Performance tuning and Configuration of Cassandra cluster.
  • Experience in development related CRUD operations in Cassandra.
  • Real time streaming the data using Spark withKafka.
  • Worked on CreatingKafkatopics, partitions, writing custom partitioner classes.
  • Application performance optimization for Cassandra cluster.
  • Monitoring the Cassandra cluster using JMX and Splunk.
  • Maintenance Cassandra cluster using Node Tool.
  • Load Balancing the Cassandra cluster.
  • Exposure on usage of ApacheKafkadevelop data pipeline of logs as a stream of messages using producers and consumers.
  • Experience in integrating ApacheKafkawith Apache Spark for real time processing.
  • Experience in development related Redshift CRUD operations.
  • Responsible for building scalable distributed data solutions usingHadoop.
  • Importing and exporting data into HDFS and Hive using Sqoop.
  • Experienced in defining job flows.
  • Processed the low latency queries.

Environment: Spark, Scala, Kafka, Cassandra, Hive, SparkSQL, MapReduce, YARN, AWS, Redshift, Java,Sqoop, Storm, JSON, XML, Eclipse, Git.

Confidential - Chicago

Hadoop Developer

Responsibilities:

  • Worked on analyzingHadoopcluster and different big data analytic tools including Pig, Hbase database and Sqoop.
  • Responsible for building scalable distributed data solutions usingHadoop.
  • Implemented nine nodes CDH3Hadoopcluster on Red hat LINUX.
  • Involved in loading data from UNIX file system to HDFS.
  • Created HBase tables to store variable data formats of PII data coming from different portfolios.
  • Implemented best income logic using Pig scripts and UDFs.
  • Writing the Hive queries using HQL.
  • Implemented test scripts to support test driven development and continuous integration.
  • Responsible to manage data coming from different sources.
  • Load and transform large sets of structured, semi structured and unstructured data.
  • Cluster coordination services through Zookeeper.
  • Designing and implementing NoSQL database stores such as MongoDB, Cassandra.
  • Experience in performance analysis and capacity planning for growing Cassandra and Hadoop clusters.
  • Experience in managing and reviewingHadooplog files.
  • Experience in creating data models and design.
  • Exported the analyzed data to the relational databases using Sqoop for visualization and to generate reports for the BI team.
  • Analyzed large amounts of data sets to determine optimal way to aggregate and report on it.
  • Supported in setting up QA environment and updating configurations for implementing scripts with Pig and Sqoop.
  • Designed and Developed full stack for Hadoop Distributed File System (HDFS) framework.
  • Including MapReduce, Hbase, Hive, Pig Framework, Zookeeper. Etc.
  • Writing Pig Latin scripts to process the data.
  • Wrote Map Reduce programs in Java to achieve the required Output.
  • Importing and exporting data into HDFS and Hive using Sqoop.
  • Written Hive queries for data analysis to meet the Business requirements.
  • Experience in managing and reviewing Hadoop log files.
  • Writing Pig and Hive UDF’s.
  • Experience in creating data models and design.
  • Experience in writing large datasets back to Cassandra.
  • Experience in development related CRUD operations.
  • Experience in Schema defining.
  • Creating Indexes and Aggregation framework.
  • Load and transform large sets of structured, semi structured and unstructured data.
  • Exported the analyzed data to the relational databases using Sqoop for visualization and to
  • Generate reports. Developed Hive queries for the analysts
  • Application performance optimization for Cassandra cluster.
  • Got good experience with NOSQL database. Involved in loading data from UNIX file system to HDFS.
  • Supported Map Reduce Programs those are running on the cluster.
  • Responsible to manage data coming from different sources.
  • Created Cassandra Advanced Data Modeling course for DataStax.
  • Supported in setting up QA environment and updating configurations for implementing scripts with Pig and Sqoop. Cluster co-ordination through Zookeeper.
  • Involved in creating Hive tables, loading with data and writing hive queries.

Environment: Hadoop, HDFS, Hive, HBase, Cassandra, Sqoop, PIG,MapReduce, Zookeeper, Java (JDK 1.6), Eclipse, PL/SQL, MySQL Shell Scripting and Ubuntu.

Confidential, Chicago

Hadoop Developer

Responsibilities:

  • Developed simple and complex MapReduce programs in Java for Data Analysis on different data formats
  • Developed workflows using Oozie to automate the tasks of loading the data into HDFS and pre-processing with Pig
  • Implemented scripts to transmit data from Oracle to HBase using Sqoop and vice-versa
  • Worked on bucketing and partitioning the HIVE table and running the scripts in parallel to reduce the run time
  • Developed and optimized Map Reduced Jobs to use HDFS efficiently by using various compression mechanisms
  • Analysed data by preforming Hive queries and running Pig scripts
  • Developed Spark Code using python for faster processing of data
  • Implemented business logic by writing Pig UDF's in Java and used various UDF's from Piggybanks and other sources
  • Worked with application teams to install operating system, Hadoop updates, patches, version upgrades as required
  • Exported the analysed data to the relational databases using Scoop for visualization and to generate reports for the BI team.
  • Implemented testing scripts to support test driven development and continuous integration

Environment: Hadoop, HDFS, Hive, Sqoop, PIG, MapReduce, Java (JDK 1.6), Eclipse, MySQL Shell Scripting and Ubuntu

Confidential, Atlanta

Java/J2EE Developer

Responsibilities:

  • Involved in the analysis, design, and development and testing phases of Software Development Life Cycle (SDLC)
  • Developed and integrated REST web services to display data or search results
  • Designed CSS based page layouts that are cross-browser compatible and standards-compliant
  • Responsible for design and development of the web pages from mock- ups
  • Designed and developed creative intuitive user interfaces that address business and end-user needs, while considering the technical, physical and temporal constraints f the users
  • Used Bootstrap library to quickly build project UI's and used AngularJS framework to associate HTML elements to models
  • Extensive experience on using Angular directives, working on attribute level, element level and class level directives
  • Utilized modular structure within the Angular JS application in which different functionalities within the application were divided into different modules
  • Developed a single page, cross-device/cross-browser web application for real-time location sharing utilizing AngularJS, JavaScript API
  • Designed dynamic and browser compatible pages using HTML5, DHTML, CSS3, jQuery and JavaScript
  • Developed code to call the web service/APIs to fetch the data and populate on the UI using JQUERY/AJAX
  • Involved in Core Java coding by using Java APIs such as Collections, Exception Handling, Generics, Enumeration, and Java I/O to fulfil the implementation of business logic
  • Participated in development of a well responsive single page application using AngularJS framework, JavaScript, and jQuery in conjunction with HTML5, CSS3 standards with front-end UI team
  • Developed front end UI using HTML5, CSS3, jQuery, JavaScript (AngularJS), AJAX and Spring for back- end development

Environment: Java, Spring MVC REST-ful, HTML5, SVN, jQuery, JavaScript, Angular JS, Oracle, Eclipse

Confidential

Java Developer

Responsibilities:

  • Generated Use case diagrams, Activity flow diagrams, Class diagrams and Object diagrams in the design phase
  • Used Java Design Patterns like DAO, Singleton etc.
  • Written complex SQL queries for retrieving and updating data
  • Involved in implementing multithreaded environment to generate messages
  • Used JDBC Connections and WebSphere Connection pool for database access
  • Used Struts tag libraries (like html, logic, tab, bean etc.) and JSTL tags in the JSP pages
  • Involved in development using Struts components - Struts-config.xml, tiles, form-beans and plug-ins in Struts architecture
  • Involved in design and implementation of document based Web Services
  • Used prepared statements and callable statements to implement batch insertions and access stored procedures
  • Involved in bug fixing and for the new enhancements
  • Responsible for handling the production issues and provided solutions
  • Configured connection pooling using WebLogic application server
  • Developed and Deployed the Application on WebLogic using ANT build.xml script
  • Developed SQL queries and stored procedures to execute the backend processes using Oracle
  • Deployed application on WebLogic Application Server and development using Eclipse

Environment: Java/ J2EE, SQL, HTML, CSS, JavaScript, JDBC, Java Beans, Eclipse, Apache Tomcat

We'd love your feedback!