Senior Big Data Engineer Resume
Hoffman Estates, IL
SUMMARY
- Senior Big Data Engineer, working on Hadoop technologies and helping clients in finding solutions leveraging big data capabilities
- Having more than 10 years of IT experience in Big data, Business Intelligence, Data warehousing, SQL Development and Data analysis
- Worked as a Technical Lead, managed client relationship for one of the top insurance industries in USA
- Involved in creating proposals for the big data project and demonstrated big data capabilities to the clients
- Strong experience in writing Map Reduce programs in Java, writing queries using Pig Latin and Hive, HBASE (NoSQL), Zookeeper, Oozie, Sqoop, Impala, Flume, Python
- Experience in developing Apache Spark applications using Python and creating real time data streaming solutions using Apache Spark Core, Spark SQL & Spark Streaming
- Proficient at using Spark APIs to cleanse, explore, aggregate, transform, and store retail data
- Having exposure to Apache Solr, Mongo DB, AWS and Google Big Query
- Worked on different data patterns like Text, JSON, XML and AVRO and used different compression techniques
- Good knowledge on object oriented programming and hands on experience using Core Java
- Experienced in working with RDMS databases Microsoft SQL, MySQL, Oracle and Teradata
- Having excellent analytical capabilities and proficient in data analysis
- Worked as a BI lead, experienced in the area of business intelligence, data warehousing, data analysis and data modeling
- Experienced working with SAP BO XI 3.1/R2/R1 and proficient in creating business interactive dashboards using Xcelsius and Tableau, developing BO Universes, and building Web - I and Crystal Reports
- Microsoft Certified Technology Specialist (MCTS) on Microsoft SQL Business Intelligence 2005 and experienced in providing BI solutions using SSIS, SSRS and SSAS
- Experienced in T-SQL programming, functions, stored procedures, views, materialized views, triggers, cursors using SQL
- Experienced in dimensional modeling, Star Schema / Snowflake Schema, Fact and Dimensional tables
- Have experience with Unix environment and Confidential scripting
- Worked both in Waterfall and Agile methodologies
- Worked in Energy & Utilities, Insurance and Retail domains
TECHNICAL SKILLS
Databases: MS SQL Server 2005/2008/2012 , Oracle 8I/9i/10g, DB2 and Teradata
ETL tools: SSIS and Ab Initio
Big data tools: Hadoop, Map Reduce, Pig Latin, Hive, Impala, Apache Spark, Spark SQL, Spark Stream, Sqoop, HBASE, Oozie, Flume, Lucene, Apache Solr, Mongo DB, AWS and Google Big Query
Development/Productivity Tools: Hue, SQL Server Management Studio, Data Transformation Services, MS SSIS, SSAS, SSRS
Hadoop frameworks: Cloudera and Hortonworks
Business Intelligence tools: Microsoft Business Intelligence and SAP Business Intelligence
Reporting tools: Business Objects Desk-I, Web-I, Crystal and MS SSRS
Dashboard tools: BO Xcelsius 2008, Tableau 8.0
Programming Languages: C, C#, Java, Python, SQL, PL/SQL, Spark SQL, HTML, Unix/ Confidential Scripts and Visual Basic
Software Engineering Tools/Technologies: SQL Server Management Studio, TOAD, Fiddler, MS Visio, VSS, TFS, OpenSVN, IBM rational ClearCase, Excel, IceScrum, Jira, Git and ServiceNow
PROFESSIONAL EXPERIENCE
Confidential, Hoffman Estates IL
Senior Big Data Engineer
Responsibilities:
- Participated in architectural discussions and designed solutions by leveraging big data capabilities
- Analyzed existing Teradata processes and prepared functional & requirements documents
- Involved in design discussions, provided optimized and cost effective solutions and prepared mapping and design documents
- Designed and implemented Apache Spark - Streaming Applications
- Used Sqoop and FastExport utilities to extract the data from Teradata
- Developing scripts using Pig Latin for the data transformations and data processing
- Built re-usable user defined functions using Java programming to utilize them in pig and hive scripts for data cleansing and data transformations
- Designed and developed Apache Spark applications using PySpark and worked with Spark SQL
- Successfully written Spark Streaming application to read streaming syw.com platform interactions messages and analyze the data
- Created hive and Impala tables on Aggregated HDFS data for analysis and this data was used to load in Google Query for Micro Strategy Reports.
- Worked on Google Cloud Platform to create Big tables and loaded data from HDFS files
- Wrote queries on Big Query for data analysis
- Involved in creating advanced full-text search engine using Apache Solr
- Created and modified collections using Mongo DB
- Created Oozie configuration and work flow to manage and schedule all the Hadoop jobs
- Involved in configuring/setting up of Apache Flume to stream the data from different APIs’
- Used Sqoop to extract the data from various different sources like Teradata and MySQL
- Create Map Reduce programs to implement custom functionalities and to use different data patterns
- Worked on the incremental process to insert, update and delete the incoming data.
- Converted the data to Avro format and used compression techniques for space utilization
- Created utility scripts using bash to standardize and automate the whole process
- Creating Hive tables with partitions to load the reference data imported from Teradata database
- Optimizing Map reduce programs to use HDFS efficiently by using various performance tuning techniques
- Working in Agile environment and experienced with IceScrum, Jira and Git tools
Environment: Cloudera framework, HDFS, Java, Map Reduce, Hive, Pig, Python, Spark, Sqoop, Oozie, Teradata, MySQL, DB2, Java (jdk 1.7), Linux, HBASE
Confidential, Oakbrook IL
Sr. Hadoop developer
Responsibilities:
- Participated in architectural discussions and provided solutions during the analysis phase of the project
- Explored various options and created POCs’ to prove the solution
- Leading a team of experienced hadoop professionals and coordinating with them for the project development
- Analyzed existing mainframe process and prepared functional & requirements documents
- Involved in design discussions, provided optimized and cost effective solutions and prepared mapping and design documents
- Currently developing hadoop applications and helping offshore team members
- Developed a custom input format to read the encrypted binary files and creating very complex map reduce programs to convert them to ASCII files for further processing
- Developing scripts using Pig Latin for the data transformations and data processing
- Built re-usable functions using Java programming to utilize them in pig scripts
- Creating Hive tables with partitions to load the reference data imported from Teradata database using Sqoop
- Logging the entire process in HBASE tables and loading the transactional into the same
- Exploring various options like Hortonworks stinger to improve the hive performance for reporting solutions
- Preparing Oozie workflows and properties files to control the business process flow and to schedule the entire process
Environment: Hadoop, HDFS, Map Reduce, Hive, Pig, Sqoop, Oozie, Teradata, DB2, Java (jdk 1.7), Linux, HBASE, Zookeeper and Tableau
Confidential, Chicago
Big Data Developer
Responsibilities:
- Worked with Hadoop team and involved in setting up the multi node cluster
- Developed a POC to create a workflow using Pig and Hive, pulled data using Sqoop and loaded into HDFS
- Understood the existing SAS model and preparing mapping and design documents
- Developed data pipelines to import/export structured/unstructured datasets using Sqoop to move data in and out of the Hadoop ecosystem
- Created Hive tables and used Pig Latin to perform transformations done in SAS model
- Created Map Reduce programs, UDF functions in Java to perform business logic
- Optimized Map/Reduce Jobs to use HDFS efficiently by using various compression mechanisms
- Used Sqoop to pull the data from Oracle database and moved the data to HDFS
- Developed workflow in Oozie to automate tasks of loading data into HDFS and pre-processing with PIG
Environment: Hadoop, HDFS, Map Reduce, Hive, Pig, Sqoop, Oozie, Oracle, Java (jdk 1.7), Linux and Hbase
Confidential, Chicago
Sr. BI Lead/Associate
Responsibilities:
- Led a team of BI analysts and managed/coordinated multiple projects
- Actively participated in discussions with the business users to understand their business priorities and proposed BI solutions like ETL, dashboards and reporting solutions
- Prepared functional requirement documents and design documents to help building the BI solutions
- Involved in preparing the conceptual documents based on the analysis
- Designed and prepared dashboard prototypes, proposed new dashboard features and gave demos to the business to show the functionality
- Involved in design discussions and prepared the design documents for ETL (SSIS & Ab Initio) and reports (BO Web-I and Crystal)
- Provided ETL solutions to gather data from different types of sources (SQL, Oracle, DB2, XML, Excel, CSV etc.) and to load into reporting layer staging tables
- Involved in designing and creating integrated data marts in Oracle by extracting data from various sources using Ab Initio and used this to used to feed strategic dashboards and reports
- Created dashboards using BO Xcelsius and Tableau for reporting
- Created BO Universes, Web-I reports, Crystal reports 2008 for the claims daily operations
- Extensively used business objects tools like desktop intelligence, Infoview, web intelligence, CMC, web-I rich client, QaaWS, Xcelsius
Environment: SAP Business Objects XI 3.1/R2/R1 (Web-I, Crystal and Xcelsius), Tableau, SQL Server 2005/2008, Oracle 10i, SSIS, Ab Initio, .Net and Web Services
Confidential
Sr. Software Engineer
Responsibilities:
- Participated in requirements gathering discussions and prepared functional and technical design documents
- Analyzed the source data model, created mapping documents and designing ETL and reporting solutions
- Worked on extracting, transforming and loading data using SSIS Import/Export Wizard and also created SSIS Packages to integrate data coming from various files (Delimited, CSV, Excel, XML etc.) and from different databases (SQL & Oracle)
- Designed and Developed multiple SSIS packages to extract data from various sources to load the data into data warehouse and provide output files to SAP systems
- Involved in creating Multidimensional cubes and designing DW schemas using SSAS 2008
- Writing multidimensional DB queries (MDX) and building and deploying SSAS cubes
- Understood the various complex reporting requirements and created report models and various reports using Reporting Services
- Worked on Charts, Matrix reports, sub reports, linked reports, drill through & drill down reports, report builder and report subscriptions
- Wrote complex T-SQL programs and prepared stored procedures, functions, cursors and views to be used as part of ETL solutions
Environment: SSIS, SSRS, MS SQL Server 2005 EE & SE(X32 and X64 bit), SQL Server 2005 DE(x64), Windows 2003\2000 Server Enterprise( x32 and X 64), MS Access, VB Script
Confidential
Software Engineer
Responsibilities:
- Involved in scope discussions with team lead and the clients
- Understood the existing Sybase code, prepared documents based on the analysis
- Involved in converting the Sybase scripts to SQL scripts
- Prepared unit test cases and performed testing on the scripts to match the results post the conversion
- Involved in deploying the converted code and provided support post production
