big Data Consultant/architect/developer/admin Resume
NC
SUMMARY:
- Over 12 years of experience in IT industry, dealing with and managing complex projects involving multiple stake holders across geographic locations.
- Hands on experience on major components in Hadoop Ecosystem like Hadoop Map Reduce, HDFS, HIVE, PIG, HBase, Sqoop, Oozie, Flume and Parquet file format (column based storage format).
- Worked extensively with Sqoop for importing and exporting the data from HDFS to Relational Database system and vice - versa. Loading data into HDFS.Involved in loading data from UNIX file system to HDFS
- Implemented Daily jobs that automate parallel tasks of loading the data into HDFS using Oozie coordinator jobs
- Experience in designing and developing applications in Spark using Scala to compare the performance of Spark with Hive and SQL/Oracle.
- Performed streaming of data into Apache ignite by setting up cache for efficient data analysis
- Analyzed the data by performing Hive queries and running Pig scripts to study customer behavior
- Well versed in installation, configuration, supporting, documentation and managing of Big Data and underlying infrastructure of Hadoop Cluster in Cloudera (CDH4 and CDH5), MapR and Hortonworks distribution
- Responsible for data analysis and design of data mapping using ETL Informatica Powercenter
- Worked extensively in Data Cleansing, Data Mining, ETL Performance Tuning of Informatica mappings and sessions
- Architected integration of various data sources/Targets with Multiple Relational Databases like Oracle and Worked on integrating data from flat files, CSV files, DB2
- Reviewed and provided detailed input to the team on ETL designs
- Ensured Master Data Management (MDM) and ETL code conforms to established coding standards and meets the feature specification
- Closely worked with DBA’s to create Physical Databases, Application Tuning etc
- Developed and maintaining Workflow Scheduling Jobs in Oozie for importing data from RDBMS to Hive
- Architected and implemented data standards, performing data analysis, business analysis, conducting Joint Application Design (JAD) sessions with Subject Matter Experts and data modelling of my assigned business areas
- Architected the Data warehouse solution using ETL (Informatica PowerCenter), PL/SQL, stored procedure and function
- Extensively worked in Oracle SQL, PL/SQL, SQLPLUS, SQL Loader, Query Performance tuning, Created DDL Scripts, Created database Objects like Tables, Indexes, Synonyms, Sequences etc., Created utility programs using UNIX shell scripts, involved in Data Import/Export
- Worked extensively on requirements management, change management and configuration management
- Experience in developing Microstrategy 10.x and Hadoop
- Experience in Self-Service BI Applications with MicroStrategy, Other Enterprise BI Tools SAP Business Objects, Tableau, Qlikview, IBM Cognos
- Ability to work independently with business partners and management to understand their needs and exceed expectations in delivering Tools/Solutions.
TECHNICAL SKILLS:
Big Data Ecosystems: Hadoop, HDFS, HBase, pig, Sqoop, Hive, Oozie, Zookeeper, Yarn, Cassandra
SPARK Streaming Technologies: SPARK, Storm
Scripting Languages: Cassandra, Scala
Programming Languages: SAS, Java, SQL, Java Scripting, HTML5
Databases: RDBMS, NoSQL, Oracle 11g, Sybase ASETools: Eclipse, JDeveloper, MS Visual Studio, Microsoft Azure HDInsight, Microsoft Hadoop cluster, JIRA.
Testing Tools: NetBeans, Eclipse.
Reporting Tools: Cognos BI, Informatica, Tableau,: SAP Business Objects (BO), Crystal Reports
Operating Systems: Unix/Linux, Windows
PROFESSIONAL EXPERIENCE:
Confidential,NC
Big Data Consultant/Architect/Developer/Admin
Responsibilities:
- Worked in sqoop to import data between RDBMS and Hadoop Distributed File System.
- Participated in Architectural Discussion for using the best practices.
- Documented the Processes and Presented them to the Architectural Team to meet the Compliances rules of the Client.
- Experience in writing Spark Applications using spark-shell, pyspark, spark-submit. Developed prototype Spark Applications using Spark-Core, Spark SQL, DataFrame APIHave Designed, Developed and Coded ETL in Hive and Pig as part of preparation of Parallel Environment to Existing Current System.
- Gathered Requirements from the existing Legacy System, then Designed and Developed a respective ETL Process in Hadoop basing on the collected requirements.
- Responsible to manage data coming from different sources and involved in HDFS maintenance and loading of structured and unstructured data.
- Developed Map Reduce jobs in Java for data cleansing, preprocessing and implemented complex data analytical algorithms.
- Developed Map Reduce programs to join data from different data sources using optimized joins by implementing bucketed joins or map joins depending on the requirement.
- Imported data from structured data source into HDFS using Sqoop incremental imports.
- Experience on SPARK streaming technologies like Kafka, Storm and created Storm data pipelines for real time processing.
- Load the dataset using SparkSQL and build tables. Get the target subset with conditions quires, and build the model.
- Created Hive tables, partitions and implemented incremental imports to perform ad-hoc queries on structured data.
- Development and Maintenance of Enterprise BI Reporting applications using MicroStrategy and Tableau.
- Work with business partners to identify areas where we can use technology to make business processes more efficient.
- Ensure that the team develops and maintains repeatable systemic processes that continue to make the team more efficient
- Consult on Reporting issues and help prioritize potential improvements
- Ensure that Technical solutions follow best practices, are reliable, are easily maintainable and are scalable under sustained loads
- Understand the existing code base to enhance the features and help fix issues in a very time critical environment
- Work with multi-relational databases and execute enhancements to applications as needed - Teradata, Oracle, Sybase ASE, SQL Server
- Identify and implement opportunities for continuous improvement regarding policies and procedures
- Participate in Agile Development Processes.
- Installed, configured and Managed Hadoop Cluster for running Big Data Applications.
Environment: s: Hadoop, Cloudera, MapReduce, Hive, SPARK SQL, Spark Streaming, Avro, Parquet Linux, Sqoop, Shell Scripting, Oozie, Cassandra, XML, Scala, Java, Oracle.
Confidential,CA
Senior Hadoop Consultant
Responsibilities:
- Developed multiple MapReduce jobs in Java for data cleansing, preprocessing and implemented CDH3 Hadoop cluster on CentOS and complex data analytical algorithms, Assisted with Performance Tuning and Monitoring.
- Experienced Spark Framework on both batch and real-time data processing
- Developed MapReduce programs to parse the raw data, populate staging tables and store the refined data in partitioned tables in the Enterprise DataWarehouse(EDW).
- Tested raw data and executed performance scripts
- Created Hive Queries that helped market analyst spot emerging trends by comparing fresh data with EDW reference tables and historical metrics.
- Created HBase tables to load large sets of structured, semi-structured and unstructured data coming from UNIX, NoSQL and a variety of portfolios.
- Created reports for the BI Team using SQOOP to export data into HDFS and HIVE.
- Managed and reviewed Hadoop Log Files.
- Expertise with MicroStrategy Architect to create re-usable schematic layer
- Experience in Big Data Hadoop Integration to MicroStrategy, Hadoop, HDFS, Pig, Hive, Oozie,
- Installed and configured Hadoop MapReduce, HDFS and HIVE.
Environment: Hadoop, Cloudera, MapReduce, Hive, SPARK, Avro, Linux, Sqoop, Shell Scripting, Oozie, Cassandra, XML, Scala, Java, Oracle.
Confidential,NY
Project Lead/Big Data Consultant
Responsibilities:- Created Hive Generic UDF's to process business logic with Hive QL.
- Optimized Hive queries, improve performance by configuring Hive Query parameters.
- Responsible for running Hadoop streaming jobs to process terabytes of XML Data.
- Development of Oozie workflow for orchestrating and scheduling the ETL process.
- Worked in retrieving transaction data from RDBMS to HDFS, get total transacted amount per user using MapReduce and save output in Hive table.
- Experience in implementing Kafka consumers and producers by extending Kafka high-level API in java and ingesting data to HDFS or HBase depending on the context.
- Worked on creating the workflow to run multiple Hive and Pig jobs to analyze very large data sets, which run independently with time and data availability.
- Developed SQL scripts using SPARK for handling different data sets and verifying the performance over Map Reduce jobs.
- Developed SPARK scripts by using Scala Shell commands as per the requirement.
- Involved in moving data from Hive tables into Cassandra for real time analytics on hive tables.
- Experienced in setting up alerts for Hadoop clusters
- Maintained Hadoop cluster which includes adding, removing cluster nodes, cluster monitoring and troubleshooting, reviewing and managing data backups and Hadoop log files.
Environment: Hadoop, Cloudera, MapReduce, Hive, SPARK, Avro, Linux, Sqoop, Shell Scripting, Oozie, Cassandra, XML, Scala, Java, Oracle.
Confidential,NY
Lead Big Data Consultant
Responsibilities:- Developed MapReduce programs to parse the raw data and store the refined data in tables.
- Designed and Modified Database tables and used HBASE Queries to insert and fetch data from tables.
- Responsible in moving all log files generated from various sources to HDFS for further processing through Flume.
- Performed loading and transforming large sets of structured, semi structured and unstructured data from relational databases into HDFS using Sqoop imports.
- Responsible for analyzing and cleansing raw data by performing Hive queries and running Pig scripts on data.
- Developed Pig Latin scripts to extract the data from web server output files to load into HDFS.
- Created Hive tables, loaded data and wrote Hive queries that run within the map.
- Used OOZIE Operational Services for batch processing and scheduling workflows dynamically.
- Hands on experience in application development using Java, RDBMS, and shell scripting.
- Performed data mining investigations to find new insights related to customers.
- Developed sentiment analysis system per particular domain using machine learning concepts by using supervised learning methodology.
- Involved in collecting the data and identifying data patterns to build trained model using Machine Learning.
- Manage and review Hadoop log files. Documented and addressed all the defects, questionable function error and inconsistencies in output.
- Attended multiple change and incident management meetings.
Environment: Java, HBase, Hadoop, HDFS, Hortonworks, Hive, Sqoop, Flume, Oozie, Zookeeper and MySQL
Confidential
Senior Big Data Consultant
Responsibilities:- Developed high-level design documents, Use case documents, detailed design documents and Unit Test Plan documents and created Use Cases, Class Diagrams and Sequence Diagrams using UML.
- Extensive involvement in database design, development, coding of stored Procedures, DDL&DML statements, functions and triggers.
- Used Sqoop to connect to the DB2 and move the pivoted data to Hive tables or Avro files
- Developed Hive queries to process the data for visualizing.
- Responsible to manage data coming from different sources.
- Worked with business teams and created Hive queries for ad hoc access.
- Loaded daily data from websites to Hadoop cluster by using Flume.
- Created complex Hive tables and executed complex Hive queries on Hive warehouse.
- Wrote MapReduce code to convert unstructured data to semi structured data.
- Used Pig to extract, transformation & load of semi structured data.
- Creating Hive tables and working on them using Hive QL.
- Documented all the changes in the system.
Environment: Hadoop, MapReduce, HDFS, Hive, Pig, HBase, Java, Cloudera Linux, XML, MySQL, MySQL Workbench, Java 6, Sybase ASE, Sybase IQ, SAP Sybase BO
Confidential
Software Engineer
Responsibilities:
- Analyze impact of new business rules, gathering requirements to introduce applications into business.
- Worked on running Update Statistics and Query Optimization and automated procedure for backing up & restoring the user databases.
- Wrote SQL, PL/SQL Stored Procedures, Functions and Packages to migrate data from SQL Server Database to Oracle Database
- Performed Database Administration of all database Objects including Tables, Clusters, Indexes, Views, Sequences, Packages and Stored Procedures.
- Implemented Oracle 11g and Upgraded existing database from Oracle 9i to Oracle 11g.
- Developed the Shell Script for generating the SQL statement for creating the Create Database statement with Alter Database Statement and setting the Database Options.
- Developed backup & recovery plan for backing up and restoring the system & user databases.
- Performance tuning of Oracle Databases & User Applications.
- Used SQL*Loader as an ETL tool to load data into staging tables
- Create Functional and Technical Design Specifications for enhancements and bug fixes.
- Performing Unit Testing, Integration Testing for the new bug fixes and enhancements on Development & Test Environments.
- Worked with QA Testing group to put the new releases on the production servers.
- Responsible for overall production support issues, overall deliverables and coordination from offshore.
- Implemented Agile, Scrum Methodologies for developing the Application.
Environment: Oracle 9i/10g/11g, Control M, Sybase ASE, PL/SQL, SQL*Loader, SQL Navigator, TOAD
Confidential
System Analyst
Responsibilities:
- Gathering high level business requirements from client and translate them into low-level tasks and then design Oracle PL/SQL code to accomplish these objectives.
- Develop complex Stored Procedures, Functions, Packages, Triggers and Oracle Database Logic.
- Worked with Development Team to improve Oracle Application Performance and resolve Real-Time problems as they occur.
- Worked in capacity of Application Developer and Oracle, Sybase Database Developer.
- Going through the Client’s Requirements and coming up with Database Design & Application Architecture.
- Monitoring and Optimizing Database Performance by monitoring Memory, CPU Utilization, Disk Utilization, Locks, Deadlocks, Runtime of Queries.
- Providing consultation on Sybase to our Development Teams.
- Designed database & developed SQL Queries and Stored Procedures, Tables using Sybase.
- Created Indexes to help optimize queries and improve performance.
- Fine-tuned SQL Queries and Stored Procedures to improve performance.
- Provided 24-7 on call production support for the Application and Database.
- Project lead for Application Development and Maintenance.
Environment: Sybase ASE, Control M, Oracle 9i/10g/11g, PL/SQL, SQL*Loader, SQL Navigator, TOAD
Confidential
System Programmer
Responsibilities:- Performed Database Administration of all database objects including Tables, Clusters, Indexes, Views, and Procedures.
- Responsible for overall deliverables and co-ordination in Team
- Involved in Logical & Physical Database Layout Design
- Set-up and Design of Backup and Recovery Strategy for various databases.
- Performance Tuning of Oracle Databases and User Applications.
- Provided User Training and Production Support.
- Improved in Performance of the Application by re-writing the SQL Queries.
- Used TOAD tool to perform Oracle related Procedure creation.
- Developed UNIX shell Scripts using KSH shell for file transfer.
Environment: Oracle 10g, Toad PL/SQL, UNIX, VB 6.0, BMC Remedy
