Big Data - Senior Consultant Resume
5.00/5 (Submit Your Rating)
SUMMARY:
- Having 10+Years of relevant experience in software development and Strong work experience in Big Data Analytics from end to end with architecture design and build a new platform for the development process
- Experience in Big Data Architecture and Data Science Model, including design in all phases of development with Hadoop Eco Systems and setup Analytical Platform, which is used by Data Scientist and Analytical team.
- Expertise in Big data technologies SPARK, HDFS, Map Reduce/YARN, Hive, PIG, Sqoop, Impala and Oozie and experience with Core Java, Scala, Python.
- Strong development skills on Map Reduce and Spark applications, experience in troubleshooting, debugging and optimizing Map Reduce with Eco systems and Spark Applications.
- Build Analytical Platform with Data warehousing on top of Hadoop, high level data summarization to RDBMS and cluster management activities with optimal cost.
- Experience in Data Warehouse design, data modeling and have knowledge on ETL and Reports
- Having 7+Years’ experience in Banking Domain and 3+Years’ experience in Financial domain, good exposure in Enterprise Risk Modelling, am working on various risk applications like Enterprise Risk Information System(ERIM), Anti Money Laundering (AML), Fraud Crimes Technology(FCT) and Credit Risk Modelling(CRP).
- A creative, energetic results - oriented engineer with broad technical and management experience in storage solutions, capacity planning, performance tuning, server management, problem management, and contingency planning and datacenter migration.
- Good interpersonal and communication skills, Collaborate with data architects, modelers and IT team members on project.
- Working on agile life cycle management in Scrum Kanban methodology and working with Jira/version story boards.
TECHNICAL SKILLS:
Operating Systems: Linux, Windows Family
Infrastructure/Project Design: Microsoft Visio, MS Office
Document store: Confluence, SharePoint, Office 365
Hadoop Eco System: Spark, Nifi, Hive, Pig, Sqoop, Map Reduce/YARN, Impala, Oozie.
Database Servers: Teradata, MS SQL Server 2005/2008, Greenplum
Data warehouse Design Tools: Teradata Load Utilities (Bteq, Fast, Multi load), SSIS, SSRS, SSAS
Programming Languages: Java, Scala, Python, Shell Scripts
PROFESSIONAL EXPERIENCE:
Confidential
BIG DATA - SENIOR CONSULTANT
- Understanding the customer 360 behavior based on data scientist and analytics manager requirement and provide the data to scoring or data modeling process.
- Working on the neural network(NN) algorithm to generate the vector data and create a design and write the code in spark to provide the data to data scientist
- Handling entire development and support area of this project and coordinate with Infrastructure team to setup the Hadoop Eco Components like spark, hive, sqoop, Greenplum etc.
- Handling a team and Design Development operation task with team members to start the development process based on the business requirement
- Create Analytics Base Table(ABT), prepare the Model Ready Data(MRD) and provide the data to modeling and scoring team. Performance analysis, tuning, developed models and execute the python script for modeling algorithm.
- Creating structured web data for fid.com site using dynamic python/shell scripts.
- Performance tuning the spark/hive scripts(job) and advice to team members to find the errors and give the perfect approach.
Confidential
BIG DATA - SENIOR ANALYST
- The CRP is the only platform to collect the data from various upstream, the technology wise it may be various databases like Teradata, Oracle, SQL Server, Netezza, DB2, MYSQL, Mainframe with file system using SPARK with Scala to load the data into Data Frame and Dataset to implement the all business functionality applying through Spark transformation and applying function’s like filter by, group by, aggregate etc.
- Risk Data is structure oriented, so we implement Spark SQL using the Data Source, Dataset, Data frame APL based on the scenario we apply all API relevant functions are using to aggregate, split, validate the columns and implement the memory management techniques based on file size,
- We convert Data frame to persist table and write SQL based on scenario, in other case we load the data into data frame applying map, reduce functions using spark with scala.
- Design the Hadoop data pipeline Framework(HDPF) and implement to source the data in various platforms, the sourcing layer have various parts like Raw, sanitization, etl, file conversion, Data Integrity.
- The HDPF is based on SPARK and fully dynamically source and write the data into HDFS. The SPARK implementation based on the Java programming language.
- The HDPF is five major activities Sourcing, Data Sanitization, Data Integration(Merge), Transformation and Publish.
- Design the Big Data Architecture and create the schema and choose the appropriate ECO system for the relevant process.
- Download the load ready file from mainframe, which is source file for the process and it’s in EBCDIC or ASCII format, download using NDM process in mainframe.
- Load the data from load ready file (Source file) to HIVE binary table, here we used both PIG and HIVE eco systems based on SOR and Convert the binary data into AVRO serde format.
- Apply the VAP or business logic which is same as in legacy, we implemented various vap’s are Collateral, Guarantor and delinquency and create a view for the business process
- Write the User Defined Functions (UDF) in java and used the UDF in PIG or HIVE
- Implementing Scrub for DELETE and UPDATE in HIVE.
- Each process we are creating workflow and design the process depends on logic, so we implemented combination of HIVE, PIG, SQOOP, SPARK and Hadoop shell commands using OOZIE.
- Setting the analytical environment named SABER, which we apply analytical validation, functions and prediction and calculate predictive analytics and discuss with the business.
- After Merging layer, the file converts to PARQUET file format and publish the data into publish layer, after that creating two types of views, one is Masked view and another one is Unmasked view and giving SENTRY access to the views and integrating with IMPALA. The end user or downstream people can access the data from Impala.
- Collect all sources like xml, Mainframe, Flat file and Teradata based on the sources design the big data architecture.
- Based on the Flat file set the properties in Hive table which is specified field terminator and collection terminator and load the data into Hive table.
- Sqoop the data from Teradata into Hive table and create the partition based source system, machine and channels.
- Create the User Defined function (UDF) and implemented in HIVE.
- Design and develop Modifying the database structure, as necessary, from information given by application developers, monitoring and optimizing the performance of the database, maintaining archived data.
- Analyzing and mining business data to identify patterns and correlations among the various data points.
- Migrate large volume of data from various sources to Database and follow the Data conversion, Data cleansing, Data validation and massaging to ensure accuracy and quality of data.
Confidential
DATA WAREHOUSE - DEVELOPER
- Creating a database Architect, ER Diagrams, Table design, Data dictionary depends on the Project Frame Work, Creating Process Data Flow diagram for each process.
- Create Complex stored procedures, which manage the entire process, Implement Performance tuning for long running queries.
- Provided the security for source files using cozyroc encryption/Decryption Standards, Dynamically create the connection managers in SSIS Package.
- Creating Scorecard, Aggregation Reports.
- Implemented the Incremental Model in the approach data can be loaded using SSIS Package.
- Recommended and supervised development of SQL Server Reporting Services as the primary of revenue, promotion, profit based Analysis Report. Implemented push and pull reporting using email, Excel, and PDF files.
- Designed data model and Dynamic stored procedures can be implemented in the project.
- Trigger mail can be implemented. It automatically send the mail to customer before crosschecking the unsubscribe process.
- Dynamically read the file (Excel, CSV, and TSV) Extracted, transformed, cleansed, Conversion the Bulk Data loaded into database.
- Dynamic stored procedures can be implemented in the project and implemented SSRS overview Report.
Confidential
DATABASE - ANALYST
- Dynamically read the file (Excel, CSV, and TSV) Extracted, transformed, cleansed, Conversion the Bulk Data loaded into database.
- Design and develop SSIS packages call from Application itself.
- Analysis of company revenue, promotion, profit based OLAP using SQL Server Analysis Service exported to Excel Pivot format.
- Recommended and supervised development of SQL Server Reporting Services as the primary of revenue, promotion, profit based Analysis Report. Implemented push and pull reporting using email, Excel, and PDF files.
- Designed data model and stored procedures to provide service-oriented database architecture.
- Operator Billing Implementation using WCF Service.
Confidential
APPLICATION DEVELOPER
- Worked as programmer and analyzed user specifications for workability, completeness and business flow.
- Developed code DAL with Business logic, create Design pattern with class implementation using C# with ASP.NET application.
- Hierarchy based payment release in parent child relationship, stored procedure using for hierarchical based Queries.
- Analyzed requirements and design for Manager Application
- Stored procedure for frequently used queries
- Analyzed requirement, Designed user interface.
- Developed code and logic for Game Progress Module, Game Schedule Module and Manage Master Data.
- Admin Control Panel to set permission for individual users.
- Automated mail with Ajax implementation for pages.
