Senior Bigdata Engineer Resume
Charlotte, NC
PROFESSIONAL SUMMARY:
- Around 9 years of experience in development, implementation and testing of Business Intelligence and Data Warehousing solutions.
- Around 4 years of experience in Big Data analytics using Hadoop, HDFS, MapReduce, Hive, Spark, Sqoop, Oozie, AWS, Nifi, Snowflake, YARN, Spark Streaming, Scala, Kafka, Flume, Pig, HBase.
- Experience in installing, configuring and administrating Hadoop cluster of major Hadoop distributions.
- Experience in installing, configuring and using Hadoop, HDFS, Hive, Pig, HBase, Sqoop and Flume.
- Experience in developing the custom UDFs for PIG and Hive
- Excellent knowledge in Hive and Pig Analytical functions
- Experience in Elastic Search, Solr and Kibana
- Experienced in different distributions of Hadoop like IBM BigInsights, Hortonworks and Cloudera.
- Experience in development, implementation and testing of Database projects.
- Experience in Data Warehousing and ETL using Informatica Power Center.
- Strong experience in Architecture, Analysis, Design, Development and Implementation of Business Intelligence solutions using Data Warehouse/Data Mart Design, ETL, OLAP.
- Extensive experience in Data Warehousing, Data Modeling using Star Schema and Snow - Flake Schema, Physical and Logical Data Modeling.
- Worked extensively with complex mappings using Expressions, Joiners, Routers, Lookups, Update strategy, Source Qualifiers, Aggregators to develop and load data into different target types.
- Strong experience in Relational Database concepts and ER diagrams.
- Working with many popular Relational Database Management Systems like IBM DB2, Oracle and MS SQL Server
- Experienced in working on Agile projects
- Experience in developing webservices using SOAP API
- Extensive experience in creating the Workflows, Worklets, Mappings, Mapplets, Reusable transformations and scheduling the Workflows and sessions using Informatica PowerCenter
- Extensive experience using Microsoft software products including Microsoft Office Suite (Word, Excel, Access, Outlook, PowerPoint and Publisher) for Windows 7/NT/XP and Vista.
- Strong conceptual, analytical, and design skills and excellent communication skills with leadership qualities.
- Excellent team work spirit and capable of learning new technologies and concepts.
- Worked under stringent deadlines with teams as well as independently.
TECHNICAL SKILLS:
Big Data Technologies: Hadoop (HDFS & MapReduce), Spark, Hive, Sqoop, AWS, Nifi, Snowflake, Flume, Zookeeper, Oozie, Spark streaming, Kafka, Flume, Elastic Search, Solr, Kiabana, HBase, Pig
Languages /Scripting: Java, C++, C, Perl, Python, PHP, Shell Scripting, HTML, XML
ETL Tools: Informatica PowerCenter (Source Analyzer, Mapping Designer, Mapplet, Transformations, Workflow Monitor, Workflow Manager)
Databases and Tools: Oracle, Teradata, Aster, SQL, PL/SQL, Toad, SQL Developer, Tableau
Platforms: Windows, Unix(Solaris), Linux(Ubuntu), VMWare
IDEs: Eclipse, Netbeans
Concepts: Data Structures
EXPERIENCE:
Confidential, Charlotte, NC
Senior BigData Engineer
Responsibilities:
- Extensively worked in business and functional requirement analysis, understanding source systems thoroughly by creating the design process flow used for standard BigData Implementation.
- Installed/Configured/Maintained Apache Hadoop clusters for application development based on the requirements.
- Developed framework for automated data ingestion from different sources like relational databases, delimited files, JSON files, XML files into HDFS and build Hive/Impala tables on top of them
- Developed spark based ingestion framework for ingesting data into HDFS, creating tables in Hive and executing complex computations and parallel data processing.
- Developed real-time data ingestion application using Flume and Kafka
- Developed data ingestion pipeline into AWS S3 buckets using Nifi
- Created external and permanent tables in Snowflake on the AWS data
- Developed java applications to read data from third party APIs and ingest data into HDFS in micro-batches
- Used Flume to read messages from ActiveMQ brokers and write the files to HDFS
- Developed Impala queries for faster querying and perform data transformations on Hive tables
- Developed application to clean semi-structured data like JSON/XML into structured files before ingesting them into HDFS
- Developed Spark code using python for Pyspark and Spark-SQL for faster testing and processing of data
- Used Hbase/Pheonix to support front end applications that retrieve data using row keys
- Developed Hive UDFs using java as per business requirements
- Used python to parse XML files and created flat files from them
- Worked on different file formats like Parquet, ORC, Avro and Compression techniques like Gzip, Snappy in Hadoop
- Automated the data ingestion using Oozie workflows and scheduled jobs using Control-M scheduler
- Used Bit Bucket as the code repository and frequently used Git commands to clone, push, pull code to name a few from the Git repository
- Used Jira as an agile tool to keep track of the stories that were worked on using the Agile methodology
- Developed application to refresh PowerBI reports using automated trigger API
- Continuous monitoring and managing the Hadoop cluster using Cloudera Manager
Environment: Cloudera CDH 5.9.16, Hive, Impala, Spark 2.4, Kafka, Flume, AWS, Nifi, Snowflake, Java, Shell-scripting, SQL, Sqoop, Oozie, Java, Oracle, SQL Server, HBase, BitBucket, Control-M, PowerBI
Confidential, McLean, VA
Senior BigData Engineer
Responsibilities:
- Extensively involved in business and functionality requirement analysis, understanding source systems thoroughly by creating the design process flow used for standard BigData Implementation.
- Extensively worked on configuration and integration of Hortonworks Hadoop platform with Solr and SOAP webservices.
- Worked on configuring connections between HDFS, Solr and Webservices.
- Worked on architectural design for the project based on BigData Enterprise Standards.
- Worked on developing store and retrieve operations in webservices.
- Imported documents into HDFS, HBase and creating HAR files.
- Imported metadata related to loan documents into Solr and provide the file path of document in HDFS
- Worked in configuring security for HDFS, Solr and Webservices.
- Worked on developing oozie jobs to create HAR files
- Monitoring the integrity of the services by developing checksum algorithms
- Worked on UDFS using Python for data cleansing
- Worked with hundreds of terabytes of data collections from different loan applications into HDFS.
- Worked in developing encryption on sensitive data at rest in HDFS, Solr and HAR files.
- Worked in data movement from one cluster to another cluster after HAR compressions for further analytics on the documents.
- Worked in versioning of the documents in HDFS and Solr.
- Worked in creating POCs for multiple business user stories using Hadoop ecosystem.
- Worked in presenting the demo of the services to different enterprise teams.
Environment: Hadoop, Hortonworks 2.2 and 2.4, HDFS, Solr, HAR, HBase, Oozie, Agile, SOAP API webservices, Java, Weblogic
Confidential, Johnscreek, GA
Hadoop / Information Architect
Responsibilities:
- Involved in installing and configuring BigInsights Hadoop platform including ecosystem environment on the server.
- Involved in installing and configuring Sysncsort.
- Involved in importing data from Mainframes to Hadoop using Sysncsort.
- Involved in creating tables, configuring permissions in BigSQL and Hive.
- Involved in configuring connections between Mainframes, Hadoop and Visualization tools like Tableau and Business Objects
- Involved in data ingestion from different RDBMS sources into Hadoop using Sqoop.
- Created POC using BigSQL, BigR, BigSheets, Text analytics
- Created POC to ingest customer clickstream realtime data using Apache Spark streaming, Scala and Kafka.
- Created SSIS jobs to import MAIN metadata from different sources like Oracle, DB2, SQL server into Oracle MAINODS database for MAINCat
- Installed and configured ElasticSearch and Kibana
- Created jobs to import data from Oracle MAINODS to ElasticSearch.
- Created Kibana dashboards to search and visualize the MAINCat data.
- Involved in POCs for CREDIT using Hortonworks and Cloudera.
- Installed and configured Hortonworks and Cloudera distributions on single node clusters for POCs
- Involved in demos for different business solutions using Cloudera and Hortonworks.
Environment: Hadoop, BigInsights, Hive, BigSQL, Sqoop, Syncsort, Oracle 11g, DB2, SSIS, Elastic Search, Kibana 4, Tableau, PL/SQL, SQL Server, SQL Developer
Confidential, Duluth, GA
Big Data/Hadoop Consultant
Responsibilities:
- Involved in installing and configuring Hadoop, Hive and Pig environment on the server
- Involved in importing data from relational databases like Teradata, Oracle, MySQL using Sqoop
- Created MapReduce jobs in processing the raw data
- Created Hive and PIG to process the raw data
- Created jobs to load data from relational databases into Hadoop using Oozie scheduler.
- Involved in importing the raw log files from different server into HDFS using Flume
- Involved in creating MapReduce jobs on log files and creating tables in HBase
- Involved in creating HiveQL on HBase tables and importing efficient work order data into Hive tables
- Involved in configuring the connection between Hive tables and reporting tools like Tableau, Excel and Business Objects
- Involved in exporting the processed data into Aster database for analysis
- Involved in creating the SQL-MR analytic functions using aster analytics
- Experience in analyzing the data using Hive and Pig
- Involved in reviewing and managing the Hadoop log files
- Involved in POC in processing of data using Spark and Kafka.
- Monitoring the Hive, Pig and Sqoop jobs on oozie scheduler
- Involved in designing reports and dashboards using Tableau.
Environment: Hadoop, Pig, Hive, Java, Sqoop, HBase, Teradata, Aster, Tableau, noSQL, Oracle 10g, PL/SQL, SQL Server, SQL Developer Toad, SuSE Linux
Confidential, Pleasanton, CA
BigData/Application Developer
Responsibilities:
- Installed and configured Apache Hadoop, Hive and Pig environment on the prototype server
- Configured MySql Database to store Hive metadata
- Responsible for loading unstructured data into Hadoop File System (HDFS)
- Created POC to store Server Log data in MongoDB to identify System Alert Metrics
- Created Reports and Dashboards of Server Alert Data
- Created Map Reduce Jobs using Pig Latin and Hive Queries
- Built Big Data Edition & Hadoop based architecture remodelling for one reporting stream.
- Data is collected from Teradata and pushing into Hadoop using Sqoop
- Used Sqoop tool to load data from RDBMS into HDFS
- Cluster coordination services through Zoo Keeper
- Automated all the jobs for pulling data from FTP server to load data into Hive tables, using Oozie workflows
- Created Reports and Dashboards using structured and unstructured data
- Maintained documentation for corporate Data Dictionary with attributes, table names and constraints.
- Extensively worked with SQL scripts to validate the pre and post data load.
- Created unit test plans, test cases and reports on various test cases for testing the data loads
- Worked on integration testing to verify load order, time window.
- Performed the Unit Testing which validate the data is processed correctly which provides a qualitative check of overall data flow up and deposited correctly into targets.
- Responsible for post production support and SME to the project.
- Involved in the System and User Acceptance Testing.
- Involved in POC working with R for data analysis.
Environment: Hadoop, Cloudera, Pig, Hive, Java, Sqoop, HBase, noSQL, Informatica Power Center 8.6, Oracle 10g, PL/SQL, SQL Server, SQL Developer Toad, Windows NT, Stored Procedures.
