Sr. Data Engineer Resume
Palo Alto, CA
SUMMARY
- Over 10 years of rich industry experience in Insurance, Finance, Banking, Distribution domains, and vast experience in Database and Big Data related applications and extensive coding experience in distinct programming languages.
- Involved in various phases of SDLC such as requirement gathering, analysis design, development, testing, UAT, deployment and maintenance.
- Deep expertise in Hadoop ecosystem (YARN(MR2), HDFS, Hive, Hue, Sqoop, Oozie, Flume, Zookeeper, Flume, Solr, Spark, Scala, Python and Kafka)
- Designed and developed applications with Spark - Scala - Hive - Solr integration.
- Hands On experience on developing SparkSql scripts to handle data transformations.
- Created various pipelines to load data from KAFKA to HBase and Hive.
- Hands on experience on running spark applications and tracking the Application, YARN and Container Logs.
- Effectively used Big data loading tools like Streamsets, Confidential Bigdata Hub, Confidential Vora.
- Created and configured AWS instances using AWS Import instance command line tools.
- Loaded third party operating system Images into AWS and spin up Ec2 instances from the image.
- Effectively made use ofTable Functions, Indexes, Table Partitioning, Collections, Analytical functions, Materialized Views, Query Re-WriteandTransportable table spaces.
- Designed and developed staging tables, Conversion routines (SQL Loader Scripts), Custom PL/SQL API's.
- CreatedTables, Views, Constraints, Index(B Tree, Bitmap and Function Based).
- Developed Complex database objects likeStored Procedures, Functions, Packages and Triggers using SQL and PL/SQL.
- Experience inOracle supplied Packages, Dynamic SQL, RecordsandPL/SQL Tables.
- Worked extensively onRef Cursor, External Tables,Collections, Dynamic SQL, CollectionsandException handling.
- Proficient in development methodologies such as Agile, Scrum and Waterfall.
- Experienced in creating complex mappings using various transformations, and developing strategies for Extraction, Transformation and Loading (ETL) mechanism.
TECHNICAL SKILLS
Programming Languages: SQL, Oracle PL/SQL, UNIX Shell Scripting, Java,J2EE- Servlets, JSP, JDBC, Java Script, mySql
Big Data Technologies: Hadoop, HDFS, AWS, Sqoop, Hive, HBase, Flume, beeline, Tableau, Hue and Zookeeper, Spark, Scala, Python and Kafka.
Special Software: HP PPM, OBIEE, Oracle 11g, SQL Server, Tableau 9.0, Informatica Power Center 9.1/ 8.6.1, Crystal Reports 10, Github, Putty, SqlDeveloper, PL/SQL Developer, TOAD, WinSQL, GitHub, JIRA and Eclipse and Pycharm.
PROFESSIONAL EXPERIENCE
Confidential, Palo Alto, CA
Sr. Data Engineer
Responsibilities:
- Developed datastrategies and analytical solutions using bigdata tools and technologies.
- Implemented solutions for ingestingdatafrom various sources utilizing Confidential BigDatatechnologies such as Confidential Bigdata hub, Confidential Vora.
- Work in cross-disciplinary teams within Confidential to understand client needs and develop applications which helps business to monetize various internal data sources.
- Utilize leading Big Data methodologies, such as Hadoop, Spark, Python, AWS, Confidential HANA, and Confidential Big Data Hub.
- Helpeddatascientists to createdatapipelines from AWS and preprocess thatdatafor modelling and machine learning purpose.
- Currently working ondatalake initiative to move on premisedatato AWS cloudDataLake and create monitoring and alerting fordataprocessing jobs.
- Involved withdatascientists to create product recommendation engine for clients using Hadoop and deployed that solution in AWS cloud.
- Analyzed and profiled data and built lake and data models.
- Manage AWS resources by creating and managing users, groups and roles and implement security policies.
- Perform data cleansing, data enrichment, data transformations, data flow and data mapping on various data sources.
- Design and develop ETL processes with stored procedures to render the models.
- Work closely with security team to adhere to Organization and Industry standards.
Confidential, San Francisco, CA
Sr.Big Data Developer
Responsibilities:
- Working on a Dataplatform to implement Bigdata solutions using Hive, MapReduce, Unix Shell scripting, and Spark and Scala technologies.
- Build
- Responsible for building scalable distributed data solutions usingSpark and Scala programing techniques.
- Developed shell scripts to invoke hive HQL scripts and created appropriate dependency.
- Developed SQOOP scripts and OOZIE workflows for scheduling the processes.
- Handled the importing of data from various data sources, performed transformations using Hive, Spark, loaded data into HDFS and extracted data from MS SQL Server into HDFS using Sqoop.
- Determined Fraud money transactions through Spark and Scala integration with SQL Server.
- Created various pipelines and workflows for data Ingestion through Big Data Hub.
- Messages are produced and consumed via KAFKA to stream to SOLR.
- Developed Parallel and Multithread processing through Par collections.
- Created various Scala Classes, Traits and functions.
- Extensively used SparkSQL to analyze Hive Tables and SOLR Collections
- Analyzed Risk data by performing Hive queries.
- Exported the analyzed data to the relational databases using Sqoop for visualization and to generate reports for the BI team.
- Responsible for Data Ingestion like Flume and Kafka.
- DevelopedSparkapplications for Batch and Real-time Processing.
- Extensively created Data Frames for various transformations and actions on incoming data.
- Integrated SOLR and Hive Data inSPARKapplications.
- Involved in installing, configuring and managing Hadoop Ecosystem components like SPARK, SCALA, HIVE, SQOOP, OOZIE, HBASE and KAFKA.
- Configured Tableau and generated the reports and dashboards using the tool Tableau.
- Worked towards continuous performance enhancements of Hive queries.
- Experience in managing and reviewingHadooplog files.
- Provide optimization recommendations and solutions for existing processes and designs.
Confidential, Foster City, CA
Big Data Developer
Responsibilities:
- Analyzed the data by performing Hive queries.
- Exported the analyzed data to the relational databases using Sqoop for visualization and to generate reports for the BI team.
- Responsible for Data Ingestion like Flume and Kafka.
- DevelopedSparkPrograms for Batch and Real Time Processing.
- DevelopedSparkStreaming applications for Real Time Processing.
- Created Data Frames handle inSPARKwith Scala.
- Configured Tableau and generated the reports and dashboards using the tool Tableau.
- Experienced on loading and transforming of large sets of structured, semi structured and unstructured data.
- Worked with Linux systems and MySQL database on a regular basis as well with NoSQL database such as Cassandra.
- Worked towards continuous performance enhancements of Hive queries.
- Experience in managing and reviewingHadooplog files.
- Very good experience with both MapReduce 1 (Job Tracker) and MapReduce 2 (YARN) setups.
- Review and analyze logs for Yarn MapReduce Jobs.
- Provide optimization recommendations and solutions for existing processes and designs.
- Designed conceptual data model based on the requirement, interacted with non-technical end users to understand the business logics.
- Modeled the logical and physical diagram of the future state, Delivered BRD and the low-level design document.
- Discussed the Data model, data flow and data mapping with the application development team.
- Developed Tableau visualizations and dashboards using Tableau Desktop.
Confidential - Mountain View, CA
Sr. Database Developer
Environment: Oracle 11g, SQL, PL/SQL, Toad 10.1, Java 1.7,J2EE- Servlets, JSP, JDBC, Jboss, Windows NT, UNIX Shell Scripting, HP PPM, Piper, Perforce.
Responsibilities:
- Involved in Integration of HP PPM with Piper and Minority Report applications.
- Designed and developed PPM request types for Minority Report application.
- Responsible for System Integration testing and migrations after UAT.
- Involved in writing queries and fetching the data from the database using PL/SQL on various modules like Demand, Time and Project.
- Created Request Types and Package workflows for Deployment and Demand Management.
- Responsible for the data refresh activities from Prod to Stage and Dev.
- Created PL/SQLstored procedures, functions and packagesfor moving the data from staging area to data mart.
- UsedBulk Collectionsforbetter performanceand easy retrieval of data, by reducing context switching between SQL and PL/SQL engines.
- Creation of database objects liketables, views, materialized views, procedures and packages using oracle tools likeToad, PL/SQL DeveloperandSQL* plus.
- Partitionedthe fact tables andmaterialized viewsto enhance the performance.
- Extensively usedbulk collectionin PL/SQL objects for improving the performing.
- Createdrecords, tables, collections(nested tables and arrays) for improving Query performance by reducingcontext switching.
- Extensively used the advanced features of PL/SQL likeRecords, Tables, Object typesand Dynamic SQL.
- Handled errors usingException Handlingextensively for the ease of debugging and displaying the error messages in the application.
- Developed JSP pages and client side validation byjavascript tags
Confidential
Sr. PL/SQL Developer
Environment: PL/SQL, Oracle 10g, Toad 8.0, SQL Developer, Informatica Power Center 9.x/8.x, SQL Server 2005, Windows NT, UNIX Shell Scripting, HP PPM, Star Team.
Responsibilities:
- Client interaction for requirements gathering and analysis.
- Developed AdvancePL/SQL packages, procedures, triggers, functions, IndexesandCollections to implement business logic usingSQL Navigator. Generated server sidePL/SQL scriptsfordata manipulationand validation and materialized views for remote instances.
- Experience inDatabase Application Development, Query Optimization, Performance Tuningand DBAsolutions and implementation experience incomplete System Development Life Cycle.
- Design and development of PPM Reports and Portlets as per user requirements.
- Provide post-implementation, application maintenance and enhancement support to the client.
- Worked onSQL*Loaderto load data from flat files obtained from various facilities every day. Used standard packages likeUTL FILE, DMBS SQL, and PL/SQLCollections and usedBULKBinding involved in writing database procedures, functions and packages for Front End Module.
- Involved in creatingUNIX shell Scripting. Defragmentation of tables, partitioning, compressing and indexes for improved performance and efficiency. Involved in table redesigning with implementation of Partitions Table and Partition Indexes to makeDatabaseFaster and easier to maintain.
- Worked on Informatica Power Center 9.1 Tool - Designer, Work Flow Manager, Work Flow Monitor and Repository Manager.
- Parsing high-level design specification to simple ETL coding and mapping standards.
- Applied slowly changing dimensions like Type 1 and 2 effectively to handle the delta Loads.
Confidential
JAVA & PL/SQL Developer
Environment: SQL, PL/SQL, Oracle 10g, HP PPM, SQL Server 2005,SQL Developer, STAR Team, Putty, Shell Scripting, Informatica Power Centre 8.6, BOXI.
Responsibilities:
- Change Request's (Enhancing Request types, Workflows, Portlets, and Reports).
- Effectively planned and executed production deployments.
- Created Request Types and Workflows from scratch.
- Process with supporting dashboards and portlets.
- Involved in writing queries and fetching the data from the database using PL/SQL on various modules like Demand, Time, Project and Finance.
- Responsible for System Integration testing and migrations after UAT.
- Developed Portlets and dashboards for trend analysis of reports on weekly basis.
- Customized, administered, trained, and supported HP PPM
- Developed custom JSP Reports.
- Extensively worked on data extraction, Transformation and loading data from various sources like Oracle, SQL Server and Flat files.
- Responsible for all activities related to the development, implementation, administration and support of ETL processes for large scale data warehouses using Informatica Power Center.
- Strong experience in Data Warehousing and ETL usingInformatica Power Center 8.6.
- Had experience in data modeling using Erwin,Star SchemaModeling, andSnowflakemodeling, FACT and Dimensions tables, physical and logical modeling.
- Hands on experience in tuning mappings, identifying and resolving performance bottlenecks in various levels like sources, targets, mappings and sessions.
Confidential
JAVA & PL/SQL Developer
Environment: SQL, PL/SQL, Oracle 9i, Informatica Power Center 8.1.1, UNIX Shell Scripting, JAVA, JSP, Crystal Reports.
Responsibilities:
- Developing Stored Procedure using PL/SQL for Report Generation.
- Developing the reports using crystal reports according to client requirements.
- Coding, Unit Testing and debugging the Reporting & Statements module.
- Perform Component Testing.
- Maintenance activity such as fixing high priority bugs.
- Creation of site layout/user interface from provided design concepts by using standard HTML/CSS practices.
- Worked in Production Support Environment as well as QA/TEST environments for projects, work orders, maintenance requests, bug fixes, enhancements, data changes, etc.
- Wrote conversion scripts usingSQL, PL/SQL, stored procedures, functionsandpackagesto migrate data from SQL server database to Oracle database.
