Principal Engineer Resume
Westford, MA
SUMMARY
- Data warehouse/Business Intelligence professional with 11 years of experience; Worked in all phases of project like Business Requirement Gathering/Analysis, Development, Testing, Implementation, Deployment, Maintenance and Production Support.
- Sound exposure to AWS architecture and services like Kinesis, S3, DynamoDB, EC2, EMR, Redshift, Data Pipeline, Athena, RDS, SNS.
- Extensive experience in Hadoop eco - system components such as Hive, MapReduce, HDFS, Sqoop, Yarn, Spark core, Spark SQL, ZooKeeper and Oozie; have in-depth knowledge of Hadoop Architecture, Map Reduce, JobTracker, TaskTracker, NameNode, Data Node, manage and external tables.
- Experience in writing programming languages like Python, Scala, Shell script and Java.
- Hands-on experience in building big data analytics solution, focusing on high-availability, fault tolerance, and auto-scaling.
- Solid understanding of Software as a Service (SaaS) model with AWS.
- Extensive experience in Travel, Customer Care, Banking, Logistics and Retail domains.
- Experience in creating Jenkins job to launch the Docker container for micro-services to perform the tests.
- Expertise in Extraction, Transformation, and Loading (ETL) data from various sources into Data Warehouse and Data Marts using Informatica Power Center 10.1/9.6/9.1/8.1/7.1 (Administrator activities, Repository Manager, Mapping Designer, Workflow Manager, Workflow Monitor, Worklet, Mapplets, Transformations, Partitions, version control and performance tuning).
- Experience in importing and exporting the data using Sqoop from HDFS to Relational Database systems and vice-versa.
- Extensive experience in performance tuning of ETL and database objects using advance techniques for faster execution.
- Expertise in scheduling BI/ETL jobs using Control-M, Cron and AutoSys.
- Sound knowledge on Data Warehouse concepts - Data Modeling using Star Schema/Snowflake Schema, Normalization, OLAP, ODS, EDW, Data Marts, Facts & Dimensions tables.
- Experience in visualization tools like MicroStrategy 8.x/9.x (MicroStrategy Desktop, Web Interface), MicroStrategy Narrowcast, QlikView and OBIEE.
- Experience in Production Support (rotational On-Call) for ETL and Reporting jobs - debugging issues and fixing within SLAs. Also involved in maintenance work, coordinating with other teams and managing services for all BI Servers.
- Experience in working closely with business product owners within the Agile/Scrum development process that allows business to make better decisions.
- Quick learner and ability to work in a team as well as individually, who likes to work on cutting edge technologies, problems and leverage opportunity to enhance skills.
TECHNICAL SKILLS
AWS Services: Kinesis, S3, EC2, EMR, Athena, Redshift, Data Pipeline, Glue, IAM and SNS
Hadoop echo-system: Hive, HDFS, Sqoop, Map Reduce, Oozie, Spark
Languages: Python, Scala, Shell Script, Java
Streaming: Kinesis, Kafka
ETL: Informatica, SQL Loader, Informatica Cloud, Dell Boomi
SQL Database: Oracle 9i/10g/11g/12c, PostGreSQL, SQL Server, Teradata, MySQL
NoSQL: Cassandra, DynamoDB
Scheduling: Control-M, Cron, Autosys
Version control: GIT, Stash, Perforce, WinCVS
OLAP: MicroStrategy, OBIEE, QlikView
Other Tools: Docker, Jenkins, SoapUI, Jira, Splunk, DataDog, SumoLogic, Slack
PROFESSIONAL EXPERIENCE
Confidential, Westford, MA
Principal Engineer
Responsibilities:
- Involved in designing the data-flow so data process time can be minimized and making sure no data loss while processing real-time streaming data.
- Created reusable http-publisher application using akka and akka-http in Scala that allow us to perform the REST API requests via HTTP protocol and route the messages to AWS Kinesis .
- Created consumer application in Scala to read the real-time data from Kinesis Data stream and apply the logic as per business rules before loading the data into Cassandra Database for real-time dashboards .
- Responsible for monitoring and troubleshooting all the applications/jobs to make sure data is available for reports/dashboards within SLA.
- Created testing framework using Scala , Jenkins and SoapUI to validate the messages to make sure applications changes have no impact on existing test cases.
- Dockerized the micro-services for any enhancements or bug fixes and run/test into EC2 testbed .
- Used Jenkins for continuous integration and build automation to create/pull the Docker images and push the image to a Docker registry in the AWS cloud.
Confidential
Sr. ETL Engineer
Responsibilities:
- Coordinating and interacting with key business users, project stakeholders for gathering Requirements and implementing the necessary ETL changes according to business requirements
- Design Cloud Architecture on AWS, created python script to Spin up the cluster to process the data with cost optimization rules.
- Created job in python to convert multiple source files having different column sets into JSON format file and post to AWS S3 for internal/external users.
- Optimized the Hive tables using optimization techniques like partitions and bucketing to provide better performance with HQL queries.
- Implemented external tables in oracle database for ETL jobs which improved the performance drastically, saving 20-30% of time in ETL load.
- Completed POC using Informatica cloud and Dell Boomi connecting to various sources like Oracle, Flat files and load the data into Snowflake computing.
Confidential
BI data Engineer
Responsibilities:
- Played major role in migrating the jobs from Hadoop eco-system to AWS cloud and performed testing to make sure data loaded as expected.
- Created jobs for QA framework to push the JSON messages to Apache Kafka topic which passes through the storm topology and loads into S3 buckets, making sure s3 data is as expected.
- Involved in creating python script to launch the EMR cluster and add steps to run the hive jobs for loading the data into S3 buckets.
- Created sqoop jobs to read data to/from data warehouse and importing/exporting them from/to HDFS.
- Implemented & maintained the branching and build/release strategies utilizing GIT and STASH.
- Developed Apache Spark jobs using Scala for faster data processing and used Spark SQL for querying.
- Created alerts in Splunk for all BI API jobs for invalid messages.
- Used Jenkins for CI/Automation tool for Continuous Integration
- Played major role in migration of Informatica jobs from 9.1 to 9.6 and oracle 11g to 12c.
- Worked closely with business product owners within the Agile/Scrum development process framework to deliver as per business requirements.
- Created QA framework for DW and AWS jobs to improve the data quality and automated the same in Control-M to run daily.
- Developed Informatica mappings to process high volume of data and build a data warehouse from various sources like Oracle, Teradata, flat files and XML files.
- Created script for API call to get the source files from FTP servers and transfer it to ETL server.
- Installed and configured Power Center 9.6 on UNIX platform. Upgraded to Informatica Power Center 9.6 from version 9.1. Installed Hotfix, utilities, and patches released from Informatica Corporation.
Confidential
BI Engineer
Responsibilities:
- Involved in building scalable distributed data solutions using Hadoop.
- Developed HiveQL scripts for performing transformation logic and load the data into HDFS.
- Ensured that all support requests are properly approved, documented, and communicated and closed within SLA. Also handled rotational on-call production support for the ETL and Reporting batch runs, and debugging issues whenever they occur.
- Created Run Book for developments/projects so team can be familiar with job process and have these documentations handy during production support.
- Created Hive Managed and External tables defined with static and dynamic partitions.
- Created shell script to clean up all the unwanted old files as per retention policy to free up the space in Servers and scheduled in cron job, this helped in term of cost and getting out of space issue from almost all BI servers.
- Created Deployment group in Informatica and Label to promote the code to another environment like QA and PROD.
- Managed users account and providing required access to BI systems such as Informatica, MicroStrategy, QlikView, Hadoop, AWS, also involved in maintenance work, coordinating with other teams and managing services for all BI Servers.
- Involved in creating dashboards/scorecards in MicroStrategy and QlikView.
- Created Visio diagram for Production mappings to understand the load process for any new user.
- Created jobs in Control-M and Cron to automate the ETL jobs as per required dependencies.
