Hadoop Consultant Resume
New York, NY
PROFESSIONAL SUMMARY:
- Having 7 years of experience in the field of information technology by managing and leading software development, data management and providing high quality solutions.
- Experience in planning, installing, configuring, maintaining, and monitoring Hadoop Clusters using Apache, Cloudera (CDH3, CDH4, CDH5) distributions.
- In depth knowledge of Job Tracker, Task Tracker, Name Node, Data Nodes and MapReduce concepts
- Hands - on experience on major components in Hadoop Ecosystem components like Hadoop MapReduce (MR), HDFS, HBase, Yarn, Oozie, Hive, Impala, Sqoop, Pig, Flume, HBase, Hue, Zookeeper.
- Experience in managing the Hadoop infrastructure with Cloudera Man a ger.
- Integrating HA Hadoop Cluster with Analytical tools and Provisioning tool like MS Azure.
- Configured Pseudo-distributed and Fully distributed Hadoop Clusters.
- Experience in importing and exporting the data using Sqoop from HDFS to Relational Database systems/ mainframe and vice-versa
- Experience in Hadoop Shell commands .
- Experience in Setting up Data Ingestion tools like Flume, Sqoop, SFTP and NDM.
- Having experience of Cloudera and Horton works distributions.
- Experienced in developing MapReduce programs using Apache Hadoop for working with Big Data.
- Experienced using Sqoop to import/export data into HDFS from RDBMS.
- Having knowledge about Hcatalog, Zookeeper, Cassandra, MongoDB and Neo4j.
- Having basic knowledge about real-time processing tools Storm, Kafka, Spark .
- Worked on different file formats like JSON, AVRO, Parquet, RC and ORC.
- Having Experience on monitoring tools Cloudera Manager and Ambari.
- Real time experience on Production deployment, code fixes.
- Experience in installation, configuration, supporting and managing - Cloud Era’s Hadoop platform along with CDH 4&5 versions.
- Experienced in managing and reviewing Hadoop log files.
- Loading log data directly into HDFS using Flume.
- Having knowledge on VMware installation and usage.
- Setup alerts with Cloudera Manager about memory and disk usage on the cluster.
- Worked on Tableau data visualization tools.
- Managing and scheduling batch Jobs on a Hadoop Cluster using Oozie.
- Strong experience in Java/ J2EE/ We b Technologies like HTML, Java Script, XML, XSD, CSS, J2EE, JDBC, Servlets, JSP, Java Beans, EJB, JNDI, JAXP, JAXB, SOAP, WSDL, Struts, and iBatis.
- Used Zookeeper to provide coordination services to the cluster.
- Written Hive queries for data analysis and to process the data for visualization.
- Strong understanding of the software life cycle methodologies Agile Scrum, Waterfall and various Maintenance, Development, Testing, Production Support and improving quality of final deliverables to meet Business Goals.
- Worked on SQL, MySQL and Oracle 10g.
- Development experience with IDE’s eclipse
- Development and deployment experience on Web logic 10.X.
- Development experience on persistence framework like iBATIS 3.x
- Designing and deploying Service Oriented Architecture (SOA) thru web services.
- Excellent ability to communicate with people who have varying levels of understanding of Application development, production support includes writing/executing test cases.
- Excellent interpersonal and communication skills, creative, research-minded, technically competent and result-oriented with problem solving and leadership skills.
- Ability to work independently to help drive solutions in fast paced/dynamic work environments.
- Strong team building, conflict management, time management and meeting management skills.
TECHNICAL SKILLS:
Big Data Ecosystems: Hadoop, HDFS, YARN, MapReduce, Hive, Pig, HBase, Zookeeper, Sqoop, Oozie, Flume, Parquet, Spark
Frameworks: JPA,J2EE, JSP, Servlets, Struts, Hibernate, .NET Framework 4.5
Methodology: Agile software development
Languages: Java, Hive QL, Pig Latin, R, Regex, Advanced PL/SQL, SQL, VBA, C++, C, Unix and Shell
Scripting Languages: HTML, CSS, JavaScript, DHTML, XML, JQuery
Web Technologies: Java, J2EE, Servlets, JSP, JDBC, XML, AJAX, SOAP, Restful
Architectures: SOA, Cloud Computing(AWS, EC2)
Application Server: Apache Tomcat, Glassfish 4.0, Web Logic
Database Systems: Apache Hadoop, Oracle 11g/10g/9i, DB2, MS-SQL Server, MySQL, MS-Access
Development Tools (IDEs/): JIRA, Clear case, Tableau, Splunk, RStudio, Eclipse/Net Beans, Toad, SQL Developer, AWK
Tools: Remedy 2.0, OEM, Attachmate Reflexion - X, Putty, E-Horizon, Export / Import, RMAN, Data Pump, Winscp3, ERWIN. VMWare, GridApp
Platforms: UNIX, Windows, Ubuntu(Linux), Mac OS
PROFESSIONAL EXPERIENCE:
Confidential, New York, NY
Hadoop Consultant
Responsibilities:
- Involved in running Hadoop jobs for processing millions of records of text data
- Data validation between existing system and new cluster.
- Written the Apache PIG scripts and Python to process the HDFS data. Developed Map Reduce program for parsing and loading into HDFS information.
- Developed PIG UDF'S for manipulating the data according to Business Requirements and also worked on developing custom PIG Loaders.
- Developed Pig Latin scripts for transformations, sort, group, event joins, filter.
- Assisted in exporting analyzed data to relational databases using Sqoop.
- Developed multiple Map Reduce jobs in java for data cleaning and preprocessing
- Involved in analysis, ETL design and development for extracting data from different interfaces like SQL, Flat Files.
- Involved in creating Hive tables, designing patterns, and loading and analyzing data using hive queries
- Involved in Design, develop Hive Data model, loading with data and writing Java UDF for Hive. Used the Used hIVE Windowing and analytical functions for data Analysis.
- Worked with file formats TEXT, AVRO, JSON files and involved in loading data from LINUX file system to HDFS.
- Performed Data Loading Techniques through Hive and HBase, ETL through Talend.
- Hands on experience in writing Spark SQL scripting.
- Analyzed business requirements and cross-verified them with functionality and features of NOSQL databases like HBase.
- Created HBase tables to store variable data formats of data coming from different applications.
- Worked on Transporting data from HBASE to HIVE using map reduce and HIVE - HBase storage handlers.
- Identifying Hadoop Configuration changes. Integrated test plan across the applications.
- Experience in moving all files generated from various sources to HDFS for further processing through Flume.
- Peer Code Reviews and assisting the team in full project life cycle implementation (requirement gathering to production). Working on Data stage to fetch various reports.
- Provided data flow after merging various workflows.
- Delivering complex PL/SQL queries to Testing team to test data flow across different system
- Providing SQL script to support and test team to capture data issues in production database
Environment: Hadoop, Apache Pig, Hive, Sqoop, MapReduce, Python, HBase, Talend, Toad Oracle, SQL PLUS, Linux.
Confidential, Schaumburg, IL
Hadoop Developer
Responsibilities:
- Helped business processes by developing, installing and configuring Hadoop ecosystem components that moved data from individual servers to HDFS.
- Installed and configured MapReduce, HIVE and the HDFS; implemented CDH3 Hadoop cluster on Centos. Assisted with performance tuning and monitoring.
- Worked in the BI team in the area of Big Data Hadoop cluster implementation and data integration in developing large-scale system software.
- Developed MapReduce programs to parse the raw data, populate staging tables and store the refined data in partitioned tables in the EDW.
- Capturing data from existing databases that provide SQL interfaces using Sqoop.
- Worked extensively with Sqoop for importing and exporting the data from HDFS to Relational Database systems/mainframe and vice-versa. Loading data into HDFS.
- Develop and maintains complex outbound notification applications that run on custom architectures, using diverse technologies including Core Java, J2EE, SOAP, XML, JMS, JBoss and Web Services.
- Created Hive queries that helped market analysts spot emerging trends by comparing fresh data with EDW reference tables and historical metrics.
- Enabled speedy reviews and first mover advantages by using Oozie to automate data loading into the Hadoop Distributed File System and PIG to pre-process the data.
- Provided design recommendations and thought leadership to sponsors/stakeholders that improved review processes and resolved technical problems.
- Shared responsibility for administration of Hadoop, Hive and Pig, Worked with Spark and Scala.
- Developed Hive queries for the analysts and Managed and reviewed Hadoop log files.
Environment: Hadoop, MapReduce, HDFS, Hive, Java, Cloudera, MapR, DataStax, IBM DataStage, PL/SQL, SQL*PLUS, Toad, Windows NT, UNIX Shell Scripting.
Confidential, Northbrook, IL
Hadoop Developer
Responsibilities:
- Handled importing of data from various data sources, performed transformations using Hive, MapReduce.
- Migrated the needed data from MySQL into HDFS using Sqoop and importing various formats of flat files in to HDFS.
- Worked on Hive queries to categorize data of different claims.
- Involved in loading data from LINUX file system to HDFS.
- Wrote Pig & Hive scripts to analyze customer data and detect user patterns.
- Written Hive UDFs to extract data from staging tables.
- Worked on UNIX shell scripts for business process and loading data from different interfaces to HDFS.
- Involved in creating Hive tables, loading with data and writing hive queries which will run internally in map reduce way.
- Developed Simple to complex Map reduce Jobs using Hive and Pig.
- Continuous monitoring and managing the Hadoop cluster by using Cloudera Manager.
- Managed and reviewed Hadoop log files.
- Developed Pig Latin scripts to extract the data from the web server output files to load into HDFS.
- Load and transform large sets of structured, semi structured and unstructured data.
- Responsible to manage data coming from different sources.
Environment: Hadoop, MapReduce, HDFS, Hive, Pig, Java, Shell Script, MySQL, Sqoop, Java, Eclipse, Hcatalog, Cloudera.
Confidential
Software Engineer
Responsibilities:
- Involved in analysis, estimation and development of the project.
- Used Struts MVC framework to enable the interactions between JSP/View layers.
- Involved in the client meeting and call individually and with team.
- Status reporting (weekly) on project progress
- Preparing release notes and taking care of all deployment activities
- Involved in XML parsing coding.
- Responsible for Build process to deploy the application.
- Involved in Creating Patches, Code Base and Sanity Code Checking’s for Code Release to Client.
- Completion of development on time, within scheduled plan.
- Doing QA for testing of requirement done.
- Did testing (Unit testing, System Integration testing and regression testing)
- Code review and given the Technical Support to the Team.
- Involved in writing test cases for unit testing, functional testing, integration testing.
- Regular client communication for requirement related clarification.
- Involved in CMMI L5 QA and project management stuffs.
Environment: J2EE (Servlets, JSP, Struts Frame work), JSTL, Search Engine(Autonomy), MySql, OpenLdap(Light weight Directory Access Protocol), AMS(Access Management System), Tomcat 5.5, Apache solr, CMS (Escenic 4.3), Eclipse IDE, SQL Query Browser, DWR, Omniture.
Confidential
Software EngineerResponsibilities:
- Involved in requirement gathering, analysis, estimation and development of the project.
- Used Spring MVC framework to enable the interactions between JSP/View layers.
- Involved in integrating Spring framework with the Volantis mobility server.
- Involved in challenging task of making the mobile website render perfectly on a wide category of mobile devices and their browsers such as Apple iPhone, Blackberry, HTC Touch Pro, Motorola RAZR V3, Nokia Symbian.
- Involved in creating layouts, themes and velocity templates for rendering the contents of the site on those various mobile devices.
- Involved in the client meeting and call individually and with team.
- Status reporting (weekly) on project progress.
- Preparing release notes and taking care of all deployment activities.
- Involved in implementing compatible image carousal for different browsers and mobile devices with the volantis server.
Environment: Java, J2EE (Servlets, JSP, Struts Frame work, Spring), Volantis mobility server 5.2, XDIME, Oracle, Castor, CMS Clickability, Apache Tomcat, Eclipse IDE, SQL Query Browser, Omniture, Windows.
