We provide IT Staff Augmentation Services!

Big Dataengineer/ Developer Resume

5.00/5 (Submit Your Rating)

Atlanta, GA

SUMMARY:

  • A hands on technologist having 11+ year of experience in building Data Products, Big Data & ETL solutions, Data Integration, Data Management, Data Warehousing & Data Modelling experience in enterprises with functional knowledge in US mortgage industry.
  • Expertise in working with Big Data &Informatica client tools - Designer (Source Analyzer, Warehouse designer, Mapping designer, Mapplet Designer, Transformation Developer), repository Manager, Workflow Manager and Workflow Monitor. strong experience designing and implementing Big Data solution architecture from ground up using cutting edge Big Data technologies.
  • Strong hands-on experience implementing big data solutions using technology stack including Hadoop MapReduce, Hive, HDFS, Spark, Scala API, Spark SQL, HBase, Sqoop, Flume and Oozie.
  • Expertise in Relational and Dimensional Modeling techniques like Star, Snowflake Schema, Fact and Dimensional Tables and Physical and Logical Data Modeling
  • Involved in Design and implementation of DataMarts and Data warehouse.
  • Created Oracle Store Procedures, Packages, Functions and Triggers.
  • Expertise in Performance Tuning of SQL queries & Views using Indexes, Partitions, Explain Plan, SQL Trace, on Oracle databases.
  • Created Views, Materialized views and used Cursors &Dynamic SQL.
  • Define and deliver framework, methodology & metrics for big data solutions development and modules required for data integration, marts and warehouse.
  • Discuss with stakeholders to understand current needs and potential areas for data solutions and strategize directions accordingly.
  • Work with global client base and strategize and develop software products
  • Develop and implement data products, data warehouse, data marts, operations data stores, data services solutions for the enterprise.
  • Review high level requirements and collaborate with product owners, designers and architects to refine epics, stories, tasks, acceptance criteria and software designs.
  • Utilize latest tools and technology to develop enterprise-wide solution and other cross-functional data as needed
  • Oversee creation and validation of software design specifications
  • Manage confidentiality of data by following established security/confidentiality standards
  • Reviewing specifications, designs, test plans and milestones goals.
  • Business acumen, understand business processes, challenges and manage stake holder expectations.
  • Passionately embrace commitment to exceptional service and delivery excellence.
  • Drive the successful delivery of software release in an Agile/Scrum environment
  • Implemented complex business rules in Informatica Power Center and also implemented slowly changing dimensions.
  • Expert in Data Validation and Data Analysis.

TECHNICAL EXPERTISE:

Big Data and Data Science Technologies: Hadoop, MapReduce, Pig, Hive, HDFS, Spark, YARN, Zookeeper, Sqoop, Flume, Oozie, Kafka, CDH5.x, Kite Morphline

ETL: Talend, Informatica Power Center 9.x/8.x, Repository Manger, Power Center Designer, Workflow Manager, Workflow Monitor

Cloud: Amazon EC2, Google Cloud

NoSQL: HBase

Programming Language: Java

Database Languages: SQL, PL/SQL

Databases: Oracle 11g, 10g, 9i, MySQL

Other Languages: Shell Scripting, XML

Operating Systems: Unix, Linux, Windows Server 2008

Other Tools/Languages: Open Modelshpere, Toad, SQL Navigator, MS Office

PROFESSIONAL EXPERIENCE

Confidential, Atlanta GA

Big DataEngineer/ Developer

Responsibilities:

  • Underwriting Dashboard, Analyst Performance Dashboard, Vendor turnaround time Reports are generated from Data Lake by sourcing data from different source databases (MySQL) and JSON.
  • Data loading in HDFS which includes borrower, underwriting, quality check, invoices and vendor data from different application sources using Sqoop and Kafka.
  • Change data capture and incremental data extract using Sqoop, Kafka which includes Invoices, Vendors, Originations data etc.
  • Implemented critical solution components using technologies including Hadoop, MapReduce, Hive, Spark, HDFS, Sqoop, and Flume.
  • Integration with an in-house developed PAAS (Platform as a service) application to provide single sign on, rules management, rules execution, workflow management and elastic search.
  • Created workflows and scheduling using Oozie to manage Hadoop jobs.
  • Worked closely with DevOps team to ensure infrastructure architecture is realized as designed.
  • Implemented data acquisition using Sqoop and Flume technologies.
  • Designed and implemented security and privacy strategy for data in HDFS. Created data compression and data compaction best practices and techniques.
  • Implemented major components as a proof of architecture design and to provide direction to the application team.
  • Identified and evaluated new big data technologies/products/tools that help fill the gap in overall enterprise architecture for future business needs.

Technologies Used: Hadoop, YARN, Spark, CDH 5.x, Kafka, MapReduce, Morphline, Hive, HDFS, Oozie, Sqoop, Flume, Impala, HBase, InetSoft, Zabbix

Confidential

Principle Engineer / Big Data Developer

Responsibilities:

  • Account payables and receivables reconciliation, Revenue, Invoice, pre billing summary, paid Invoices reports are generated from data hub by sourcing data from different source databases (MySQL) and JSON.
  • Data loading in HDFS which includes borrower, invoices, vendor data from different application sources using Sqoop.
  • Developed map reduce code to handle specific data inputs and for data cleaning and pre-processing.
  • Involved in import data in Hive tables using Sqoop, loading and analyzing data using hive queries and improved the Hive queries performance by implementing partitioning and clustering based on different term levels.
  • Importing and exporting of data into Hive using Sqoop and loading incremental data from MySQL databases through Sqoop and placed in HDFS and processed.
  • Created workflows and scheduling using Oozie to manage Hadoop jobs.
  • Hadoop security implementation to provide complete security for authentication, https enables using key store.
  • Map reduce unit test cases and mock classes for unit testing.
  • Implemented big data pipeline from external data sources using Sqoop and Flume technologies.
  • Designed and implemented security model for Hadoop cluster.
  • Identified and evaluated new big data technologies/products/tools that help fill the gap in overall enterprise architecture for future business needs.

Technologies Used: Hadoop, CDH, MapReduce, Hive, HDFS, Oozie, Sqoop, Flume

Confidential

Technical Leader / LeadETL Developer

Responsibilities:

  • Built a data staging, enterprise wide operational data store and a query layer
  • Built the data model by normalizing data in ODS for enterprise data needs
  • Worked on complete SDLC from Extraction, Transformation and Loading of data using Informatica.
  • Made adaptive changes to the Data warehouse environment due to changes in the source system, Business Requirements, Enhancements and Product Upgrades.
  • Worked extensively with complex mappings using Expressions, Aggregators, Filter, Joiner, Lookup, Update Strategy, Router, Rank, Sequence generator, Transaction control and Store procedures transformations to develop and feed data warehouse.
  • Wrote UNIX Shell scripts to automate workflows.
  • Utilized Informatica tasks such as: Session, Command, Timer, Email, Event-Raise and Event-Wait.
  • Setting up automated scheduling of sessions at high loads at the regular intervals.
  • Responsible for performance tuning of Informatica mappings.
  • Implemented Mapping Variables and Parameters in Transformations to calculate billing data.
  • Assisted in design and implementing of the normalized operational data store
  • Wrote Custom SQL for some complex bookmarkqueries.
  • Improving the Performance of queries.
  • Conducted meetings with the End Users for gathering and analyzing the User’s business requirements.
  • Define and document the reporting requirements based of off the business requirements.
  • Acted as a liaison between end-user groups and technical team.

Technology Used: Informatica 9.1, Oracle 11g, Toad, Shell Scripting, Linux, Windows

Confidential

Senior Software Engineer / Senior ETL Developer

Responsibilities:

  • Investor reporting system, responsible for designing and implementing the data mart. Used Open Modelsphere to design and demonstrate it to the technical team.
  • Involved in the full life cycle development of the project including Analysis, design, development, and testing.
  • Loaded data from OLTP Tables (which are in Oracle) to Staging tables (ORACLE database) and Staging tables to OLAP Tables (ORACLE tables) using Informatica 9.1.
  • Used various transformations like Source qualifier, Aggregator, Expression, Router, Filter, Update strategy, Connected and Unconnected lookup, Sorter and Router in Informatica 9.1
  • Created Mapplets using Mapplet Designer and used those Mapplets for reusable business process in development.
  • Implemented Slowly Changing Dimensions Type1 and Type2 for accessing the full history of accounts and transaction information.
  • Involved in writing the UNIX shell scripts to perform pre-session and post-session operations.
  • Developed Oracle PL/SQL programs for backend processing and exception handling
  • Developed PL/SQL code to implement business rules through cursors, ref cursors, procedures, functions, and packages.
  • Modified and developed database triggers, cursors, procedures, functions and packages to meet business requirements.
  • Identified the bottlenecks in the PL/SQL packages and improved the performance to reduce the execution time to minutes from hours.
  • Responsible for writing the technical documents and maintaining the documentation.
  • Assisted in design and implementing of the Data mart for reporting.
  • Worked closely with Business Analyst and the end users in writing the functional specifications for Data mart tables based on the business requirement needs.

Technology Used: Informatica 9.1, Oracle 11g, Toad, MS Office, Windows XP

Data Integration between various applications

Confidential

Responsibilities:

  • Built data integration/messaging platform to communicate between various applications using Perl and PL/SQL.
  • Built a near real time operational data store for different applications to consume data.
  • Data sourced from various source databases like Oracle, Sybase, MySQL, SQL Server etc.
  • Built an additional data staging layer for clients reporting data needs.
  • Responsible for the design, development and implementation of data store in efforts inCompleteSoftware Development Life Cycle using Perl and Informatica ETL tool.
  • Extensively worked with different transformations such as Aggregator, Expression, Router, Filter, Stored Procedure, and Sorter using the Informatica Power Center.
  • Used debugger to identify the issues in the data flows and fixed the mappings.
  • Implemented Slowly Changing Dimensions Type-1 and Type-2 in ETL jobs.
  • Used mapping Parameters for incremental data extraction.
  • Worked on SQL optimization. Identified bottlenecks and performance tuned the Informatica mappings/sessions.
  • Data validation of certain complicated reports. Installed and used Toad to validate the results.
  • Documented the reports that were created, the purpose and specification of the reports. Outlined the entire work on paper.
  • Used PL/SQL Store Procedures to upsert (update or insert) data for incremental tables
  • Converted existing PL/SQL Packages to ETL Mappings using Informatica.

Technology Used: Perl, PL/SQL, XML, Informatica, Oracle 9i, Toad

We'd love your feedback!