Big Data Engineer Resume
Greenville, SC
SUMMARY:
- 10+ years’ experience in Analysis, Design, Development and Implementation of Data - Warehouse (DW) and Business Intelligence (BI) Solutions.
- Over two years of experience in Hadoop Eco system and Big-Data Analytics.
- Excellent understanding/knowledge of Hadoop ecosystem and various components such as HDFS, Job Tracker, Task Tracker, Name Node, Data Node and Map Reduce programming.
- Good understanding of installing and configuring ecosystem components like Hadoop MapReduce, HDFS, HBase, Oozie, Sqoop, Flume, Pig & Hive.
- Experience in analyzing data using Pig Latin, HQL, HBase and custom MapReduce programs in Java.
- Extending Hive and Pig core functionality by writing custom UDFs.
- Experience in importing and exporting data using Sqoop from HDFS to Relational DB systems.
- Good understanding of NoSQL Databases like HBase.
- Experience in collecting business requirements, writing functional requirements and test cases and creating technical design documents.
- Excellent communication skills, interpersonal skills, problem solving skills, and a very good team player along with can do attitude and ability to effectively communicate with all levels of the organization such as technical, management and customers.
- Ability to perform at a high level, meet deadlines, adaptable to changing priorities.
TECHNICAL SKILLS:
- RDBMS: Oracle, SQL Server, DB2, ODBC, JDBC, Netezza, Teradata.
- Languages: Python, Java, SQL, PL/SQL, HTML, XML, VB.
- Hadoop 2x, including HDFS, Map Reduce, CHD 3.0 / 4.0, Sqoop, Oozie, HBase, Hive and PIG.
- Big Data Tools - Tableau - Good Understanding and POC Experience.
- Multi-tier web-based and client/server architectures.
- Linux, UNIX, MS Windows, IIS.
- MS Suite - Word, Excel, Access, Power Point, Project Plan.
- ETL/BI Tools: Informatica Power Center (Repository Manager, Designer, Workflow Monitor, Workflow Manager), Cognos (Report Studio, Query Studio, Framework Manager).
- Tools: TOAD, VSS, MS Share Point.
- Concepts and Technologies - Data Warehousing, Big Data, MDM, Planning and Forecasting, Data Mining, Data Modeling, Data Integration, Data Conversion, OLAP, OLTP, Business Intelligence.
- Oracle EBS, Oracle Hyperion DRM.
PROFESSIONAL EXPERIENCE:
Confidential, Greenville, SC
Big Data Engineer
Responsibilities:
- Gathering business process from Subject Matter Experts and understand the current system.
- Data Acquisition and sourcing for Alerts Project for Sprint 4 and Sprint 10.
- Active involvement in Architectural, Technical and process decision making.
- Lead efforts to capture, document, communicate and present all information pertinent to an end-to-end business solution including architecture and supporting recommendation.
- Converted extremely formatted historical data in excel to SQL Server using python.
- Worked towards converting current MS Access database to SQL Server following DB best practices.
- Developed data pipeline using Flume, Sqoop, Pig to ingest customer behavioural data and purchase histories into HDFS for analysis.
- Importing data from relational data stores toHadoopusing Sqoop.
- Performed joins, group by and other operations using PIG scripts for ETL and Python for (UDF).
- Worked with HBASE NOSQL database.
- Worked on Installing 20 node UATHadoopcluster.
- Helped and directed testing team to get up to speed onHadoopData testing.
- Developed SQL Server DB triggers and Stored Procs for Alerts Project in Sprint 4 and Sprint 10.
Environment: SQL Server, Python, MS Access, Business Agility, Pycharm IDE, Hadoop, Sqoop, Flume, PIG, HBase.
Confidential - Marlborough, MA
Analytics Engineer
Responsibilities:
- Gathering business requirements from Subject Matter Experts and understand the current system.
- Responsible for building scalable distributed data solutions using Hadoop.
- Analyzed data using Hadoop components Hive and Pig.
- Worked hands on with ETL process.
- Load and transform large sets of structured and semi structured using Hadoop/Big Data.
- Involved in loading data from UNIX file system to HDFS.
- Working towards Design and development of current Oracle Central Data Warehouse using HIVE
- Responsible for creating Hive tables, loading data and writing hive queries.
- Used Hive to analyze the partitioned and bucketed data and compute various metrics for reporting.
- Handled importing data from various data sources, performed transformations using Pig, Map Reduce, and loaded data into HDFS.
- Extracted the data from Oracle into HDFS using the Sqoop.
- Provided quick response to ad hoc client requests for data and reports.
- Understanding of Informatica Power Exchange to load the data from RDBMS tables to HDFS as part of a sample POC.
- Generated sample BI reports using Tableau by connecting to Cloudera Connector for HIVE QL for a POC.
- Unload reference data from DB2 to Flat files and FTP the same to HDFS.
Environment: CDH, HDFS, Map Reduce, Hive, Pig, Sqoop, Oracle, Linux, Tableau, DB2, Java, Eclipse, Cognos, Informatica Power Exchange.
Confidential, Waltham, MA
Sr. Data Analyst
Responsibilities:
- 1.5 months of extensive training to learn and understand Hadoop eco system components, advantages and future of Hadoop.
- Part of the Architecture team to identify how HDFS file system and Hadoop eco system components can be used on already existing Risk Manager Application.
- Conducted an internal POC to show Confidential Clients the advantages of using Hadoop ecosystem components and how the scalability and faster processing can be achieved.
- Implementation of the Product Platform as well as all data transfer, storage and Processing from DDW and flat files to Hadoop File Systems.
- Developed data pipeline using Sqoop to ingest data into HDFS for analysis.
- Used Pig as ETL tool to do transformations, filters, pre-aggregations, change data capture and delta record processing before storing the data onto HDFS.
- Used Hive to analyze the partitioned and bucketed data and compute various metrics for reporting.
- Developed Pig and HIVE UDFs.
- Wrote Custom Map Reduce Scripts for Data cleansing and Processing in Java.
- Importing and exporting data into HDFS using Sqoop.
- Experienced in running Hadoop streaming jobs to process terabytes of xml format data.
- Involved in developing Shell scripts to orchestrate execution of all other scripts (Pig, Hive, and MapReduce) and move the data files within and outside of HDFS.
Environment: Hadoop, Map Reduce, Hive, Pig, HBase, Java (JDK 1.6), Eclipse, CDH 3.0, Sqoop, Flume, Cognos, SQL Server SSIS.
Population Manager
Responsibilities:
- Part of data architecture team in creating ETL Mapping documents.
- Developing Logical/Physical Data-Models and Designing the DW Facts/Dimensions/Staging tables and create DDL/DML based on the model.
- Developing the ETL Load Plan and work with the ETL and Cognos Report Developers to optimize and design the reports and ETL Processes.
- Implementing Customized Care measure logic for the client by Data mining algorithms. The care measures are already defined by the Health care Norms and Confidential customizes it according to the client’s need.
- Designing PQRS (Physician Quality Reporting System) where physicians CAP is determined. PQRS has certain CAP defined for Physicians. Once a physician meets the CAP he is rewarded.
- Use and customize third party ETL tools like Rhapsody, Quality Spectrum Insight, EBM, ETG, Intelliflex to load files into Data warehouse.
- Trouble shooting, correcting and creating JIRA tickets for any defects arising in the refresh.
Risk Manager
Responsibilities:
- Used ETL (SSIS) to develop jobs for extracting, cleaning, transforming and loading data into data warehouse according to the client requirements.
- Created dynamic and customized packages for ETL from various data sources (flat files) using SSIS.
- Work with data architect in creating ETL Mapping documents for new client implementations.
- Once the data is processed in the DDW data mining algorithms are run on clients data to identify key business measures for BI Reporting.
- Framework Modeling, cubes and dashboards development in Cognos 10.
- Use Microsoft Project to schedule the project process.
- Improve the Multi-Dimensional Model with various business requirements and create fact and dimension tables.
- Constant requirement changes, field changes and new requirements coming in from the existing clients and internal product enhancements all the time which needs careful functional and regressive development.
Environment: MS SQL Server SSIS, Cognos, Business Objects xcelsius to Cognos Migration, intelliflex, Ultra Edit, Sales Force - defect tracking, JIRA and Confluence, MS SQL Server 2008
Confidential - Foster City, CA
BI Lead
Responsibilities:
- Responsibilities include leading the effort to establish new business intelligence process to develop analytical reports on the clearance and settlement data by utilizing Cognos BI tool-set.
- Served as technology and BI subject matter expert for Client Centric Solutions and research teams.Collaborated with the ETL and data architects and contributed to the overall Data Warehousing strategy.Managing five offshore BI developers to insure on-time and high quality deliverables.Documenting business requirements, providing support and training to research teams to enhance skills in using basic and intermediate DW and BI capabilities. Administered user access. Managed various in-house projects including development of brokerage applications to administer client portfolios, market-to-market and VISA products reporting. Training end-users to generate reports using Query Studio. Worked on Cognos SDK to alter the report functionalities according to the organizational requirements
Environment: Cognos 8.2 (Framework Manager, Report Studio, Analysis Studio, Query Studio), Cognos Series 7 Impromptu (IWR and Upfront), Ab Initio, Metadata Migration Utility (impcat2xml, migratetoc8), Teradata, SQL Server 2000, Windows XP.
Confidential - NY
Lead DWH Developer
Responsibilities:
- Responsibilities include working with Data Architecture teams to design logical and physical models, ETL tool evaluation and selection. Analyzing source to target mappings, validations and business rules.
- Leading team of ETL developers and also creating mappings, performance tuning informatica sessions by increasing block size, data cache size and sequence buffer length. Performed Unit and Integration Testing of Informatica Sessions, Batches and Target Data.Gathering BI business requirements and working with BA teams to develop Analytics FRD’s.
Environment: Informatica Power Center 8.1/7.1, Cognos Report Net, Windows XP, Oracle 8i/9i, HTML/XML, SQL Server2000, Confidential - DB2, Autosys, UNIX.
Confidential - Delaware
Sr. ETL&BI Developer
Responsibilities:
- Corporate Data Warehouse: Responsibilities include Understanding the business process and developing end-to-end ETL processes to extract transform and load data. Designing mappings using Informatica Source analyzer, Warehouse designer, Transformation designer and Mapplet designer. Creating and monitored tasks, workflows and sessions, using Workflow Manager and Workflow Monitor, to move data at intervals as per requirements.
- Performance tunings of mappings, sessions for optimum performance. Mentoring offshore resources. Unit and integration testing.
- Corner Stone Reporting: Responsibilities include building the conceptual Model to meet the current business requirementsand developing TO-BE Architecture in COGNOS. Designing the reporting Framework and leading report development at offshore.
Environment: Informatica 7.1, UNIX, Windows 2000, Oracle, SQL Server, Toad, Erwin, Control-M, UNIX Shell Scripting, SQL/PL/SQL, Cognos Report Net.
Confidential, NJ
ETL&BI Developer
Responsibilities:
- Developing subject oriented Data Marts using Informatica power center 6.2. ETL data from multiple input sources to Oracle. Created different transformations and used mapping designer, repository manager, server manager.
- Developing business Intelligence Reports and cubes using Cognos 7.3 Impromptu administrator, power play transformer, Impromptu Web Reports and Power Play Enterprise Server.
Environment: Cognos Impromptu 7.3, Power Play Transformer, Power Play Enterprise Server (PPES), IWR, DB2, Informatica Power Center 6.2, Toad and Erwin.
Confidential
Software Engineer
Responsibilities:
- Responsibilities include Requirement Gathering, Development in Java, HTML, and Data Model Design with the team, JDBC for Data Retrieval, SVN, and Unit Testing and Reports Development using Crystal Reports.
