Sr. Hadoop/bigdata Consultant Resume
Sacramento, CA
SUMMARY
- 8+ years of IT experience along with 4+ years of Hadoop experience.
- Experienced Big Data Engineer with different Hadoop Distributed File Systems and Eco System MapReduce, Pig, Hive, Spark, Scala, Sqoop, Oozie
- Good understanding of Hadoop Distributed File System and MapReduce.
- Experienced in installing, configuring Hadoop multi node cluster in Amazon EC2.
- Work experience in different layers of Hadoop Framework - Storage layer (HDFS), Analysis Layer (Pig and Hive), Engineering Layer (Jobs and Workflows)
- Expertise in deployment of Hadoop, Yarn,Sparkand Storm integration with Cassandra and Kafka etc.
- Experience in analyzing data using Pig Latin and Hive QL.
- Knowledge in job/workflow scheduling and monitoring tools like Oozie & Zookeeper.
- Importing & exporting data from existing databases that provide SQL interfaces using ETL tool Sqoop.
- Strong in developing, executing and scheduling Pig UDF’s for data processing
- Experienced in working with different data sources like Flat files, Spreadsheet files, log files and Databases.
- Experience on Spark and Scala.
- Developed analytical components using Scala, Spark.
- Implemented POC to migrate map reduce jobs into Spark RDD transformations using Scala.
- Developed Spark applications using Scala for easyHadooptransitions.
- Developed custom aggregate functions using SparkSQL and performed interactive querying.
- Implemented Spark using Scala and SparkSQL for faster testing and processing of data.
- Experience in monitoring and managing Hadoop cluster using Hortonworks
- Expertise with managing and reviewing Hadoop log files.
- Background with traditional databases such as Oracle, SQL Server, MySQL.
- Expertise with NoSQL databases such as HBase.
- Experience in Elastic search technologies in creating custom Lucene/Solr Query components.
- Well experienced and possess strong knowledge in Unix Shell Scripting
- Ability to work independently and in a group with effective communication and quantitative skills.
- Expertise in implementing Spark and Scala application using higher order functions for both batch and interactive analysis requirement
- Good working experience in using Spark SQL to manipulate Data Frames.
- Experiencein working with Spark tools like RDD transformations, spark SQL.
- Proficiency in formulating a strategic vision and a tactical roadmap to address client's critical Business
- Intelligence/ Analytics needs in conformance with overall corporate objectives.
- Technical evangelist skilled at developing new applications on Hadoop according to business needs, andconvert existing applications to Hadoop environment.
TECHNICAL SKILLS
Hadoop/Big Data: CDH4.4, HDFS, MapReduce2, Hive, Pig, HBase Zookeeper, Sqoop, Oozie, Flume, Storm, Spark and Scala
No-SQL Databases: Hbase
Programming Languages: Scala, Java, Pig Latin, HiveQL, Unix shell scripts, SQL, PLSQL.
Operating System: Windows, Unix, Linux, AIX 5.3/7.1
Relational Database: Oracle … DB2, SQL Server 2008, SQL Server 2012, MySQL
Tools: and IDE: Eclipse, NetBeans, IntelliJ
PROFESSIONAL EXPERIENCE
Confidential, Sacramento, CA
Sr. Hadoop/Bigdata Consultant
Responsibilities:
- Writing Scala User-Defined Functions (UDFs) to solve the business requirements.
- Creating the Case Classes.
- Working with the Data Frames and RDD’s.
- Parsing the JSONObjects into flatten formats using Scala.
- Creating the tables in Hive and integrating data between Hive &Spark.
- Worked on the core and Spark SQL modules of Spark extensively using programming languages likeScala.
- Hands on experience with Spark Scala programming and good understanding of its 'In Memory' processing capability.
- Worked on creating the RDD's, DF's for the required input data and performed the data transformations using Spark Scala.
- Experience in Kafka Producer, Consumer, Brokers.
- Experienced with batch processing of data sources using Apache Spark.
- Experienced in working with RDDs.
- Experience in Writing the Scala functions, procedures, Constructors and Traits.
- Maintenance and Monitoring of Cassandra cluster using Opscenter and Node tool.
- Cassandra data modeling and Design.
- Performance tuning and Configuration of Cassandra cluster.
- Experience in development related CRUD operations in Cassandra.
- Real time streaming the data using Spark withKafka.
- Worked on CreatingKafkatopics, partitions, writing custom partitioner classes.
- Application performance optimization for Cassandra cluster.
- Monitoring the Cassandra cluster using JMX and Splunk.
- Maintenance Cassandra cluster using Node Tool.
- Load Balancing the Cassandra cluster.
- Exposure on usage of ApacheKafkadevelop data pipeline of logs as a stream of messages using producers and consumers.
- Experience in integrating ApacheKafkawith Apache Spark for real time processing.
- Experience in development related Redshift CRUD operations.
- Responsible for building scalable distributed data solutions usingHadoop.
- Importing and exporting data into HDFS and Hive using Sqoop.
- Experienced in defining job flows.
- Processed the low latency queries.
Environment: Spark, Scala, Kafka, Cassandra, Hive, SparkSQL, MapReduce, YARN, AWS, Redshift, Java,Sqoop, Storm, JSON, XML, Eclipse, Git.
Confidential - Chicago
Hadoop Developer
Responsibilities:
- Worked on analyzingHadoopcluster and different big data analytic tools including Pig, Hbase database and Sqoop.
- Responsible for building scalable distributed data solutions usingHadoop.
- Implemented nine nodes CDH3Hadoopcluster on Red hat LINUX.
- Involved in loading data from UNIX file system to HDFS.
- Created HBase tables to store variable data formats of PII data coming from different portfolios.
- Implemented best income logic using Pig scripts and UDFs.
- Writing the Hive queries using HQL.
- Implemented test scripts to support test driven development and continuous integration.
- Responsible to manage data coming from different sources.
- Load and transform large sets of structured, semi structured and unstructured data.
- Cluster coordination services through Zookeeper.
- Designing and implementing NoSQL database stores such as MongoDB, Cassandra.
- Experience in performance analysis and capacity planning for growing Cassandra and Hadoop clusters.
- Experience in managing and reviewingHadooplog files.
- Experience in creating data models and design.
- Exported the analyzed data to the relational databases using Sqoop for visualization and to generate reports for the BI team.
- Analyzed large amounts of data sets to determine optimal way to aggregate and report on it.
- Supported in setting up QA environment and updating configurations for implementing scripts with Pig and Sqoop.
- Designed and Developed full stack for Hadoop Distributed File System (HDFS) framework.
- Including MapReduce, Hbase, Hive, Pig Framework, Zookeeper. Etc.
- Writing Pig Latin scripts to process the data.
- Wrote Map Reduce programs in Java to achieve the required Output.
- Importing and exporting data into HDFS and Hive using Sqoop.
- Written Hive queries for data analysis to meet the Business requirements.
- Experience in managing and reviewing Hadoop log files.
- Writing Pig and Hive UDF’s.
- Experience in creating data models and design.
- Experience in writing large datasets back to Cassandra.
- Experience in development related CRUD operations.
- Experience in Schema defining.
- Creating Indexes and Aggregation framework.
- Load and transform large sets of structured, semi structured and unstructured data.
- Exported the analyzed data to the relational databases using Sqoop for visualization and to
- Generate reports. Developed Hive queries for the analysts
- Application performance optimization for Cassandra cluster.
- Got good experience with NOSQL database. Involved in loading data from UNIX file system to HDFS.
- Supported Map Reduce Programs those are running on the cluster.
- Responsible to manage data coming from different sources.
- Created Cassandra Advanced Data Modeling course for DataStax.
- Supported in setting up QA environment and updating configurations for implementing scripts with Pig and Sqoop. Cluster co-ordination through Zookeeper.
- Involved in creating Hive tables, loading with data and writing hive queries.
Environment: Hadoop, HDFS, Hive, HBase, Cassandra, Sqoop, PIG,MapReduce, Zookeeper, Java (JDK 1.6), Eclipse, PL/SQL, MySQL Shell Scripting and Ubuntu.
Confidential, Chicago
Hadoop Developer
Responsibilities:
- Developed simple and complex MapReduce programs in Java for Data Analysis on different data formats
- Developed workflows using Oozie to automate the tasks of loading the data into HDFS and pre-processing with Pig
- Implemented scripts to transmit data from Oracle to HBase using Sqoop and vice-versa
- Worked on bucketing and partitioning the HIVE table and running the scripts in parallel to reduce the run time
- Developed and optimized Map Reduced Jobs to use HDFS efficiently by using various compression mechanisms
- Analysed data by preforming Hive queries and running Pig scripts
- Developed Spark Code using python for faster processing of data
- Implemented business logic by writing Pig UDF's in Java and used various UDF's from Piggybanks and other sources
- Worked with application teams to install operating system, Hadoop updates, patches, version upgrades as required
- Exported the analysed data to the relational databases using Scoop for visualization and to generate reports for the BI team.
- Implemented testing scripts to support test driven development and continuous integration
Environment: Hadoop, HDFS, Hive, Sqoop, PIG, MapReduce, Java (JDK 1.6), Eclipse, MySQL Shell Scripting and Ubuntu
Confidential, Atlanta
Java/J2EE Developer
Responsibilities:
- Involved in the analysis, design, and development and testing phases of Software Development Life Cycle (SDLC)
- Developed and integrated REST web services to display data or search results
- Designed CSS based page layouts that are cross-browser compatible and standards-compliant
- Responsible for design and development of the web pages from mock- ups
- Designed and developed creative intuitive user interfaces that address business and end-user needs, while considering the technical, physical and temporal constraints f the users
- Used Bootstrap library to quickly build project UI's and used AngularJS framework to associate HTML elements to models
- Extensive experience on using Angular directives, working on attribute level, element level and class level directives
- Utilized modular structure within the Angular JS application in which different functionalities within the application were divided into different modules
- Developed a single page, cross-device/cross-browser web application for real-time location sharing utilizing AngularJS, JavaScript API
- Designed dynamic and browser compatible pages using HTML5, DHTML, CSS3, jQuery and JavaScript
- Developed code to call the web service/APIs to fetch the data and populate on the UI using JQUERY/AJAX
- Involved in Core Java coding by using Java APIs such as Collections, Exception Handling, Generics, Enumeration, and Java I/O to fulfil the implementation of business logic
- Participated in development of a well responsive single page application using AngularJS framework, JavaScript, and jQuery in conjunction with HTML5, CSS3 standards with front-end UI team
- Developed front end UI using HTML5, CSS3, jQuery, JavaScript (AngularJS), AJAX and Spring for back- end development
Environment: Java, Spring MVC REST-ful, HTML5, SVN, jQuery, JavaScript, Angular JS, Oracle, Eclipse
Confidential
Java Developer
Responsibilities:
- Generated Use case diagrams, Activity flow diagrams, Class diagrams and Object diagrams in the design phase
- Used Java Design Patterns like DAO, Singleton etc.
- Written complex SQL queries for retrieving and updating data
- Involved in implementing multithreaded environment to generate messages
- Used JDBC Connections and WebSphere Connection pool for database access
- Used Struts tag libraries (like html, logic, tab, bean etc.) and JSTL tags in the JSP pages
- Involved in development using Struts components - Struts-config.xml, tiles, form-beans and plug-ins in Struts architecture
- Involved in design and implementation of document based Web Services
- Used prepared statements and callable statements to implement batch insertions and access stored procedures
- Involved in bug fixing and for the new enhancements
- Responsible for handling the production issues and provided solutions
- Configured connection pooling using WebLogic application server
- Developed and Deployed the Application on WebLogic using ANT build.xml script
- Developed SQL queries and stored procedures to execute the backend processes using Oracle
- Deployed application on WebLogic Application Server and development using Eclipse
Environment: Java/ J2EE, SQL, HTML, CSS, JavaScript, JDBC, Java Beans, Eclipse, Apache Tomcat
