We provide IT Staff Augmentation Services!

Senior Bigdata Engineer Resume

5.00/5 (Submit Your Rating)

Charlotte, NC

PROFESSIONAL SUMMARY:

  • Around 9 years of experience in development, implementation and testing of Business Intelligence and Data Warehousing solutions.
  • Around 4 years of experience in Big Data analytics using Hadoop, HDFS, MapReduce, Hive, Spark, Sqoop, Oozie, AWS, Nifi, Snowflake, YARN, Spark Streaming, Scala, Kafka, Flume, Pig, HBase.
  • Experience in installing, configuring and administrating Hadoop cluster of major Hadoop distributions.
  • Experience in installing, configuring and using Hadoop, HDFS, Hive, Pig, HBase, Sqoop and Flume.
  • Experience in developing the custom UDFs for PIG and Hive
  • Excellent knowledge in Hive and Pig Analytical functions
  • Experience in Elastic Search, Solr and Kibana
  • Experienced in different distributions of Hadoop like IBM BigInsights, Hortonworks and Cloudera.
  • Experience in development, implementation and testing of Database projects.
  • Experience in Data Warehousing and ETL using Informatica Power Center.
  • Strong experience in Architecture, Analysis, Design, Development and Implementation of Business Intelligence solutions using Data Warehouse/Data Mart Design, ETL, OLAP.
  • Extensive experience in Data Warehousing, Data Modeling using Star Schema and Snow - Flake Schema, Physical and Logical Data Modeling.
  • Worked extensively with complex mappings using Expressions, Joiners, Routers, Lookups, Update strategy, Source Qualifiers, Aggregators to develop and load data into different target types.
  • Strong experience in Relational Database concepts and ER diagrams.
  • Working with many popular Relational Database Management Systems like IBM DB2, Oracle and MS SQL Server
  • Experienced in working on Agile projects
  • Experience in developing webservices using SOAP API
  • Extensive experience in creating the Workflows, Worklets, Mappings, Mapplets, Reusable transformations and scheduling the Workflows and sessions using Informatica PowerCenter
  • Extensive experience using Microsoft software products including Microsoft Office Suite (Word, Excel, Access, Outlook, PowerPoint and Publisher) for Windows 7/NT/XP and Vista.
  • Strong conceptual, analytical, and design skills and excellent communication skills with leadership qualities.
  • Excellent team work spirit and capable of learning new technologies and concepts.
  • Worked under stringent deadlines with teams as well as independently.

TECHNICAL SKILLS:

Big Data Technologies: Hadoop (HDFS & MapReduce), Spark, Hive, Sqoop, AWS, Nifi, Snowflake, Flume, Zookeeper, Oozie, Spark streaming, Kafka, Flume, Elastic Search, Solr, Kiabana, HBase, Pig

Languages /Scripting: Java, C++, C, Perl, Python, PHP, Shell Scripting, HTML, XML

ETL Tools: Informatica PowerCenter (Source Analyzer, Mapping Designer, Mapplet, Transformations, Workflow Monitor, Workflow Manager)

Databases and Tools: Oracle, Teradata, Aster, SQL, PL/SQL, Toad, SQL Developer, Tableau

Platforms: Windows, Unix(Solaris), Linux(Ubuntu), VMWare

IDEs: Eclipse, Netbeans

Concepts: Data Structures

EXPERIENCE:

Confidential, Charlotte, NC

Senior BigData Engineer

Responsibilities:

  • Extensively worked in business and functional requirement analysis, understanding source systems thoroughly by creating the design process flow used for standard BigData Implementation.
  • Installed/Configured/Maintained Apache Hadoop clusters for application development based on the requirements.
  • Developed framework for automated data ingestion from different sources like relational databases, delimited files, JSON files, XML files into HDFS and build Hive/Impala tables on top of them
  • Developed spark based ingestion framework for ingesting data into HDFS, creating tables in Hive and executing complex computations and parallel data processing.
  • Developed real-time data ingestion application using Flume and Kafka
  • Developed data ingestion pipeline into AWS S3 buckets using Nifi
  • Created external and permanent tables in Snowflake on the AWS data
  • Developed java applications to read data from third party APIs and ingest data into HDFS in micro-batches
  • Used Flume to read messages from ActiveMQ brokers and write the files to HDFS
  • Developed Impala queries for faster querying and perform data transformations on Hive tables
  • Developed application to clean semi-structured data like JSON/XML into structured files before ingesting them into HDFS
  • Developed Spark code using python for Pyspark and Spark-SQL for faster testing and processing of data
  • Used Hbase/Pheonix to support front end applications that retrieve data using row keys
  • Developed Hive UDFs using java as per business requirements
  • Used python to parse XML files and created flat files from them
  • Worked on different file formats like Parquet, ORC, Avro and Compression techniques like Gzip, Snappy in Hadoop
  • Automated the data ingestion using Oozie workflows and scheduled jobs using Control-M scheduler
  • Used Bit Bucket as the code repository and frequently used Git commands to clone, push, pull code to name a few from the Git repository
  • Used Jira as an agile tool to keep track of the stories that were worked on using the Agile methodology
  • Developed application to refresh PowerBI reports using automated trigger API
  • Continuous monitoring and managing the Hadoop cluster using Cloudera Manager

Environment: Cloudera CDH 5.9.16, Hive, Impala, Spark 2.4, Kafka, Flume, AWS, Nifi, Snowflake, Java, Shell-scripting, SQL, Sqoop, Oozie, Java, Oracle, SQL Server, HBase, BitBucket, Control-M, PowerBI

Confidential, McLean, VA

Senior BigData Engineer

Responsibilities:

  • Extensively involved in business and functionality requirement analysis, understanding source systems thoroughly by creating the design process flow used for standard BigData Implementation.
  • Extensively worked on configuration and integration of Hortonworks Hadoop platform with Solr and SOAP webservices.
  • Worked on configuring connections between HDFS, Solr and Webservices.
  • Worked on architectural design for the project based on BigData Enterprise Standards.
  • Worked on developing store and retrieve operations in webservices.
  • Imported documents into HDFS, HBase and creating HAR files.
  • Imported metadata related to loan documents into Solr and provide the file path of document in HDFS
  • Worked in configuring security for HDFS, Solr and Webservices.
  • Worked on developing oozie jobs to create HAR files
  • Monitoring the integrity of the services by developing checksum algorithms
  • Worked on UDFS using Python for data cleansing
  • Worked with hundreds of terabytes of data collections from different loan applications into HDFS.
  • Worked in developing encryption on sensitive data at rest in HDFS, Solr and HAR files.
  • Worked in data movement from one cluster to another cluster after HAR compressions for further analytics on the documents.
  • Worked in versioning of the documents in HDFS and Solr.
  • Worked in creating POCs for multiple business user stories using Hadoop ecosystem.
  • Worked in presenting the demo of the services to different enterprise teams.

Environment: Hadoop, Hortonworks 2.2 and 2.4, HDFS, Solr, HAR, HBase, Oozie, Agile, SOAP API webservices, Java, Weblogic

Confidential, Johnscreek, GA

Hadoop / Information Architect

Responsibilities:

  • Involved in installing and configuring BigInsights Hadoop platform including ecosystem environment on the server.
  • Involved in installing and configuring Sysncsort.
  • Involved in importing data from Mainframes to Hadoop using Sysncsort.
  • Involved in creating tables, configuring permissions in BigSQL and Hive.
  • Involved in configuring connections between Mainframes, Hadoop and Visualization tools like Tableau and Business Objects
  • Involved in data ingestion from different RDBMS sources into Hadoop using Sqoop.
  • Created POC using BigSQL, BigR, BigSheets, Text analytics
  • Created POC to ingest customer clickstream realtime data using Apache Spark streaming, Scala and Kafka.
  • Created SSIS jobs to import MAIN metadata from different sources like Oracle, DB2, SQL server into Oracle MAINODS database for MAINCat
  • Installed and configured ElasticSearch and Kibana
  • Created jobs to import data from Oracle MAINODS to ElasticSearch.
  • Created Kibana dashboards to search and visualize the MAINCat data.
  • Involved in POCs for CREDIT using Hortonworks and Cloudera.
  • Installed and configured Hortonworks and Cloudera distributions on single node clusters for POCs
  • Involved in demos for different business solutions using Cloudera and Hortonworks.

Environment: Hadoop, BigInsights, Hive, BigSQL, Sqoop, Syncsort, Oracle 11g, DB2, SSIS, Elastic Search, Kibana 4, Tableau, PL/SQL, SQL Server, SQL Developer

Confidential, Duluth, GA

Big Data/Hadoop Consultant

Responsibilities:

  • Involved in installing and configuring Hadoop, Hive and Pig environment on the server
  • Involved in importing data from relational databases like Teradata, Oracle, MySQL using Sqoop
  • Created MapReduce jobs in processing the raw data
  • Created Hive and PIG to process the raw data
  • Created jobs to load data from relational databases into Hadoop using Oozie scheduler.
  • Involved in importing the raw log files from different server into HDFS using Flume
  • Involved in creating MapReduce jobs on log files and creating tables in HBase
  • Involved in creating HiveQL on HBase tables and importing efficient work order data into Hive tables
  • Involved in configuring the connection between Hive tables and reporting tools like Tableau, Excel and Business Objects
  • Involved in exporting the processed data into Aster database for analysis
  • Involved in creating the SQL-MR analytic functions using aster analytics
  • Experience in analyzing the data using Hive and Pig
  • Involved in reviewing and managing the Hadoop log files
  • Involved in POC in processing of data using Spark and Kafka.
  • Monitoring the Hive, Pig and Sqoop jobs on oozie scheduler
  • Involved in designing reports and dashboards using Tableau.

Environment: Hadoop, Pig, Hive, Java, Sqoop, HBase, Teradata, Aster, Tableau, noSQL, Oracle 10g, PL/SQL, SQL Server, SQL Developer Toad, SuSE Linux

Confidential, Pleasanton, CA

BigData/Application Developer

Responsibilities:

  • Installed and configured Apache Hadoop, Hive and Pig environment on the prototype server
  • Configured MySql Database to store Hive metadata
  • Responsible for loading unstructured data into Hadoop File System (HDFS)
  • Created POC to store Server Log data in MongoDB to identify System Alert Metrics
  • Created Reports and Dashboards of Server Alert Data
  • Created Map Reduce Jobs using Pig Latin and Hive Queries
  • Built Big Data Edition & Hadoop based architecture remodelling for one reporting stream.
  • Data is collected from Teradata and pushing into Hadoop using Sqoop
  • Used Sqoop tool to load data from RDBMS into HDFS
  • Cluster coordination services through Zoo Keeper
  • Automated all the jobs for pulling data from FTP server to load data into Hive tables, using Oozie workflows
  • Created Reports and Dashboards using structured and unstructured data
  • Maintained documentation for corporate Data Dictionary with attributes, table names and constraints.
  • Extensively worked with SQL scripts to validate the pre and post data load.
  • Created unit test plans, test cases and reports on various test cases for testing the data loads
  • Worked on integration testing to verify load order, time window.
  • Performed the Unit Testing which validate the data is processed correctly which provides a qualitative check of overall data flow up and deposited correctly into targets.
  • Responsible for post production support and SME to the project.
  • Involved in the System and User Acceptance Testing.
  • Involved in POC working with R for data analysis.

Environment: Hadoop, Cloudera, Pig, Hive, Java, Sqoop, HBase, noSQL, Informatica Power Center 8.6, Oracle 10g, PL/SQL, SQL Server, SQL Developer Toad, Windows NT, Stored Procedures.

We'd love your feedback!