We provide IT Staff Augmentation Services!

Data Analyst Resume

0/5 (Submit Your Rating)

Deerfield Beach, FL

SUMMARY

  • Over 8+ years of professional experience in Information Technology, which includes experience in the development of Big Data applications and extensive experience in Database development and data analysis
  • Good understanding of distributed systems, HDFS architecture, Internal working details of MapReduce and Spark processing frameworks.
  • Having good knowledge in writing MapReduce jobs through Pig, Hive, and Sqoop.
  • Extensive knowledge in writing Hadoop jobs for data analysis as per the business requirements using Hive and worked on HiveQL queries for required data extraction, join operations, writing custom UDF's as required and having good experience in optimizing Hive Queries.
  • Developed custom Kafka producer and consumer for different publishing and subscribing to Kafka topics.
  • Good working experience on Spark (spark streaming, spark SQL) with Scala and Kafka. Worked on reading multiple data formats on HDFS using Scala.
  • Extensive knowledge in writing Hadoop jobs for data analysis as per the business requirements using Hive and worked on HiveQL queries for required data extraction, join operations, writing custom UDF's as required and having good experience in optimizing Hive Queries.
  • Excellent experience in designing and developing Enterprise Applications for J2EE platform using Servlets,JSP, Struts, Spring, Hibernate and Web services.
  • Expertise in using major components of Hadoop ecosystem components like HDFS, YARN, MapReduce, Hive, Impala, Pig, Sqoop, HBase, Spark, Spark SQL, Kafka, Spark Streaming, Flume, Oozie, Zookeeper, Hue.
  • Having excellent communication, organizational and presentation skills, inter - personal relationship and ability to be part of team.
  • Liaison between the cross-functional teams, top management clients, users and any stakeholder involved in the project.
  • Dealing with processes, governance, policies, standards and tools as part of MDM to consistently define and manage the critical data of the organization.
  • Experience in creating Business Process Models for business from conceptual to procedural level i.e Designing a Project and Defining the Scope, managing the project until delivery.
  • Experienced in creating business process using use case, UML Class diagrams
  • Skilled at gathering requirements by conducting one on one interviews, brainstorming sessions and Joint Application Development (JAD) sessions.
  • Extensively involved in MDM to help the organization with strategic decision making and process improvements. (Streamline data sharing among personnel and departments).
  • Experienced in preparing process documents like Project plan (Scheduling and maintain the project timeline (Project planning/ tracking)), change request tracking, Project status report, Project issues and risk mitigation plan / reports.
  • Experience in sql and nosql
  • Good working knowledge of Amazon Web Services(AWS) Cloud Platform which includes services likeEC2,S3,VPC,ELB, IAM, DynamoDB, Cloud Front, Cloud Watch, Route 53, Elastic Beanstalk (EBS), Auto Scaling, Security Groups, EC2 Container Service (ECS), Code Commit, Code Pipeline, Code Build, Code Deploy,DynamoDB, Auto Scaling, Security Groups, Red shift, CloudWatch, CloudFormation, CloudTrail, Ops Works, Kinesis, IAM, SQS, SNS, SES.
  • Experience in Data Analysis, Data Profiling, Data Integration, Migration, Data governance and Metadata Management, Master Data Management and Configuration Management.
  • Experience in developing customizedUDF’sin Python to extend Hive and Pig Latin functionality.
  • Expertise in designing complex Mappings and have expertise in performance tuning and slowly changing Dimension Tables and Fact tables
  • Good working knowledge on data visualization and dashboard designing using Power BI.
  • Extensively usedInformaticaclient tools - Repository Manager, Mapping Designer, Workflow Manager, Workflow Monitor.
  • Strong understanding of the Analytical tools and their implementation.
  • Experience in Data Analysis, Data Profiling, Data Integration, Migration, Data governance and Metadata Management, Master Data Management and Configuration Management.
  • Strong experience in writing scripts usingPythonAPI, PySpark API and Spark API for analysing the data.
  • Experience in Google Cloud components, Google container builders and GCP client libraries and cloud SDK’s
  • Hands On experience on Spark Core, Spark SQL, Spark Streaming and creating the Data Frames handle in SPARK with Scala.
  • Experience in NoSQL databases like HBase, Cassandra and worked on table row key design and to load and retrieve data for real time data processing and performance improvements based on data access patterns.
  • Experienced in building Automation Regressing Scripts for validation of ETL process between multiple databases like Oracle, SQL Server, Hive, and Mongo DB usingPython.
  • Experience with operating systems: Linux, RedHat, and UNIX.
  • Experience as Analyst processes, interprets and documents business processes, products, services and software through analysis of data.
  • Experience as analystreviews data to identify key insights into a business's customers and ways the data can be used to solve problems.
  • Hands on learning with different ETL tools to get data in shape where it could be connected to Tableau using Tableau Data Extract.
  • Responsible for data engineering functions including, but not limited to: data extract, transformation, loading, integration in support of enterprise data infrastructures - data warehouse, operational data stores and master data management.
  • Experienced in building Automation Regressing Scripts for validation of ETL process between multiple databases like Oracle, SQL Server, Hive, and Mongo DB using Python.
  • Worked extensively on building real time data pipelines using Kafka for streaming data ingestion and Spark Streaming for real time consumption and processing.
  • Strong experience working with Hive for performing various data analysis.
  • Experienced in requirement analysis, application development, application migration and maintenance using Software Development Lifecycle (SDLC) and Python/Java technologies.
  • Extensively used Networking & Protocols TCP/IP, Telnet, HTTP, HTTPS, FTP, SNMP, LDAP, DNS.
  • Data Processing Experience in Designing and Implementing Data Mart applications, mainly Transformation Process using Informatica.
  • Extensive experience in implementation of Data Cleanup Procedures, Transformation, Scripts, Stored Procedures and execution of Test plans for loading the data successfully into Targets.

TECHNICAL SKILLS

BigData/Hadoop Technologies: MapReduce, Spark, SparkSQL,Azure,Spark Streaming, Kafka,PySpark,, Pig, Hive,HBase, Flume, Yarn, Oozie, Zookeeper, Hue, Ambari Server

Languages: C, C++, XML,R/R Studio, SAS Enterprise Guide, SAS, R (Caret, Weka, ggplot), Perl, MATLAB, Mathematica, FORTRAN, DTD, Schemas, Json, Ajax, Java, Scala, Python (NumPy, SciPy, Pandas, Gensim, Keras ), Java Script, Shell Scripting, oracle, Teradata sql, ase isql

NO SQL Databases: Cassandra, HBase, MongoDB, MariaDB

Development Tools: Microsoft SQL Studio, IntelliJ,Azure Databricks, Eclipse, NetBeans, snowflake,Teradata,postgresql

Public Cloud: EC2, IAM, S3, Autoscaling, CloudWatch, Route53, EMR, RedShift

Development Methodologies: Agile/Scrum, UML, Design Patterns, Waterfall

Build Tools: Jenkins, Toad, SQL Loader,PostgreSql, Talend,Maven, ANT, RTC, RSA, Control-M, Hue, SOAP UI,rtc,rsa,control-m,oziee,hue,Erwin,agile

Reporting Tools: MS Office (Word/Excel/Power Point/ Visio/Outlook), Crystal reports XI, SSRS, cognos.

Databases: Microsoft SQL Server, MySQL Oracle 10g/11g, 12c, DB2, Teradata, Netezza

Data Modeling Tools: Erwin r7.1/7.2, ER Studio V8.0.1 and Oracle Designer

Capital marketing tools: Equity and derivatives

Operating Systems: All versions of Windows, UNIX, LINUX, Macintosh HD, Sun Solaris

PROFESSIONAL EXPERIENCE

Confidential, Deerfield Beach, FL

Data analyst

Responsibilities:

  • Collaborated with Business Analysts, SMEacross departments to gather business requirements, and identify workable items for further development.
  • Transforming business problems into Big Data solutions and define Big Data strategy and Roadmap.
  • Installing, configuring and maintaining Data Pipelines.
  • Additional responsibility to perform POC for new tools and evaluate adaptability to the client IT framework
  • Planned, designed and implemented sql server database code objects such as tables, views and stored procedures .
  • Created sql databases tables with various constraints including primary key, foreign key, est.
  • Loaded data from flat data files into sql server database tables using table import
  • Created etl packages using sql server ssis designer to extract data from flat files and excel files
  • Huge knowledge of risk management processes.
  • Used equity tool for Accumulated other comprehensive income / loss and Retained earnings.
  • Used derivatives tool for selling stocks for the company and also predicting future and swaps between the assets
  • Initiated root cause analysis of SAP issues at the plant level as a PPPI
  • Improved plant reporting processes with 85% accuracy by identifying system gaps
  • Developed and implemented training aids for SAP and LEDS Reporting data
  • Trained Production Team Leads on processes that involve daily production reporting
  • Trained Team Leads on new PPPI processes that are implemented in SAP
  • Executed data analysis and data visualization on survey data using tableau desktop with analysis python and panda
  • Managed large datasets using panda data frames and mysql
  • Performed data cleaning, features engineering using pandas and numpy packages in python learning algorithm
  • Defined accountability procedures governing data access, processing, storage, retention, reporting and auditing measuring contract compliance
  • Assist with managing all facets of project life cycle, including design, development, testing, and deployment.
  • Developed the Stone Soup Pharmaceutical Data Model. Assisted on the development of a proposal for a major client.
  • Led development and documentation of data governance policy and procedures as well as Stewardship inventory and control set up processes for MDM projects
  • Collaborated with software architects to ensure alignment of the Netezza environment.
  • Implemented metadata standards, data governance and stewardship, master data management, ETL, ODS, data warehouse, data marts, reporting, dashboard, analytics, segmentation, and predictive modelling
  • Work with users to identify the most appropriate source of record and profile the data required for sales and service.
  • Performed project requirements gathering, requirement analysis, design and development
  • Involved in end to end deliverlabries like tables and migration etl to prod
  • Worked on various partion methods like round robin, modulus, merge, join, copyetc
  • Solely responsible for preparing release document, transition plan to support team
  • Created jobs in data stage to extract from heterogeneous data sources like oracle, sql server and flat files
  • Developed job sequencer with proper job dependencies, job control stages and trigerrs.
  • Authoring Python (PySpark) Scripts for custom UDF’s for Row/ Column manipulations, merges, aggregations, stacking, data labeling and for all Cleaning and conforming tasks.
  • Selected and generated data into csv files and stored them into AWS S3 by using AWS EC2 and then structured and stored in AWS Redshift.
  • Participated in the governance of company data and policies
  • Managed relationships between Executive Board members and members of the data governance committee
  • Scheduled business data governance meetings to establish ongoing data consistency
  • Responsible for importing data from Postgres to HDFS, HIVE using SQOOP tool.
  • Developed server based web traffic using restful api analysis tool using panda
  • Used panda api to put the data as time series and tabular format for data manipulation and
  • Architect and design serverless application CI/CD by using AWS Serverless (Lamda) application model.
  • Gathering data and business requirements from end users and management. Designed and built data solutions to migrate existing source data in Teradata and DB2 to Big Query (Google Cloud Platform).
  • Optimizing the performance of dashboards and workbooks in Tableau desktop and server.
  • Extensive Knowledge and hands-on experience implementing PaaS, IaaS, SaaS style delivery models inside the Enterprise (Data centre) and in Public Clouds using like AWS, Google Cloud, and Kubernetes etc.
  • Provided Best Practice document for Docker, Jenkins, Puppet and GIT
  • Expertise in implementing DevOps culture through CI/CD tools like Repos, Code Deploy, Code Pipeline, GitHub. rganizes different data elements and standardizes how they relate to one another and real-world entity properties. worked with Database Administrators, Business Analysts and Content Developers to conduct design reviews and validate the developed models
  • Identified, wformulated and documented detailed business rules and Use Cases based on requirements analysis
  • Using Hadoop on Cloud service (Qubole) to process data in AWS S3 buckets
  • Defined job work flows as per their dependencies inOozie.
  • Maintaining architectural principles and coding standards across the code and project lifecycles, including validating whether the data usage is as per the data security principles and that the data manipulation is built as per the requirements.
  • Created Python scripts to generate reports, insightful dashboards, market analysis and provided predictive modeling support to GTM team focused on improving performance of personalization campaigns.
  • Integrated data sources from Kafka (Producer and Consumer API) for data stream-processing in Spark using AWS Network. responsible for defining the naming standards for data warehouse
  • Built a data lake as a cloud based solution in AWS using Apache Spark and provide visualization of the ETL orchestration using CDAP tool.
  • Handled importing data from various data sources, performed transformations using Hive, MapReduce, and loaded data into HDFS.
  • Migrated existing on - premises application to AWS. Used AWS services like EC2 and S3 for small data sets processing and storage. Designed and developed Spark jobs to process batch jobs that were run on AWS EMR.
  • Imported data from AWS S3 into Spark Dataframes using Python. Performed transformations and actions on dataframes storing intermediate results in parquet format in HDFS on EMR cluster and final results on S3. Python modules were used for logging, data manipulations, Config parsing, reading environment variables etc.
  • Created data governance and privacy policies
  • Creating and modified existing data ingestion pipelines usingKafkaandSqoopto ingest the database tables and streaming data into HDFS for analysis.
  • Built real-time streaming data pipelines withKafka, Spark streamingandCassandra.
  • Created a Kafka broker in structured streaming to get structured data by schema using case classes. Configured Kafka broker for theKafkacluster of the project and streamed the data to Spark for structured streaming to get structured data by schema.
  • A continuous integration and deployment pipeline by using Jenkins and Chef.
  • Developed highly complexPythonandScalacode, which is maintainable, easy to use, and satisfies application requirements, data processing and analytics using inbuilt libraries.
  • Involved in designing optimizing Spark SQL queries, Data frames, import data from Data sources, perform transformations, perform read/write operations, save the results to output directory into HDFS/AWS S3.

Environment: Hadoop, HDFS, Kafka, Hbase, AWS,HIVE,Sqoop, Python, Spark, Sqoop, Hive, MapReduce, Docker, Hbase, Zookeeper, Oozie, Scala, java, Oracle 12c, NO SQL-Cassandra, Mongo DB

Confidential, Branch Burg, NJ

Data analyst

Responsibilities:

  • Designed, executed and optimized digital marketing and campaign on google adwords, led to 20 percent increase in roi
  • Managed redevelopment of internal tracking system in use by 125 emloyees, resulting in 20+ new features and 15 percent of operation time
  • As a Data analyst, my role includes analyzing and evaluating the business rules, data sources, data volume and come up with estimation, planning and execution plan to ensure architecture meets the business requirements
  • Maintained security and data integrity of the database
  • Designed ldm using ibm infosphere data architect data modelling and oracle
  • Schedules buissness goverance data
  • Extensively involved in MDM to help the organization with strategic decision making and process improvements. (streamline data sharing among personnel and departments)
  • Delivered consistent daily results as the fixed income capital markets office manager reporting directly to SVP and VPs of trading desk
  • Broad technology background and comprehensive exposure to various capital markets research and electronic trading platforms such as BondDesk, Valubond, TradeWeb and Bloomberg.
  • Utilized Informatica toolset (Informatica Data Explorer, and Informatica Data Quality) to analyze legacy data for data profiling. support and development of medical research information systems for data acquisition and management; Database-driven Python applications; Migration, Data transformation and validation tool development; Technical lead for several projects.
  • Evaluated data profiling, cleansing, integration and extraction tools (e.g. Informatica).
  • Coordinate with the business users in providing appropriate, effective and efficient way to design the new reporting needs based on the user with the existing functionality.
  • Support and development of medical research information systems for data acquisition and management; Database-driven Python applications; Migration, Data transformation and validation tool developmentRemain knowledgeable in all areas of business operations in order to identify systems needs and requirements
  • Wrote and ran SQL and Impala, Hadoop queries and databases including MySQL.
  • Loading Data from Curted layer to Azure database for use of border uses with Data Factory.
  • Ran and wrote dozens of queries in Impala to debug Hadoop.
  • Used portals to troubleshoot Docker and Kubernetes applications.
  • Using Kafka and integrating with the Spark Streaming.
  • Developed data pipeline using Sqoop, HQL, Spark and Kafka to ingest Enterprise message delivery data into HDFS.
  • Developed business plans, determine Key Performance Indicators (KPI) and coordinate the measurement result.
  • In charge of the ISO 9001:2015 Quality Management System biannual internal and external audits.
  • Conduct business process analysis and identify critical issues and gaps for an established organizational process.
  • Create and implement simple statistical models such as logistic regression or forecast modeling to identify risks.
  • Responsible for maintaining inventory of marketing materials for the financial planners and in the firm reception area.Responsible for keeping the schedules for all of the planners in the firm.
  • Assisted in developing new marketing materials for the firm.
  • Responsible for maintaining the marketing materials in the front waiting area.
  • Developed highly optimized Spark applications to perform various data cleansing, validation, transformation and summarization activities according to the requirements.
  • Used iam tools like oia and oim to load the feeds with user access information
  • Implemented end user requirements by creating Stage Variables and coding the business rule logic for transformations and rejections with a Business Rule Staging
  • Designed jobs to extract, cleanse and parameterize the jobs to allow probability during the running and to apply rules and logic at transformation stage
  • Developed and implemented job sequences to integrate the date stage jobs
  • Designed and developed nd delivered projects within time by all usaa quality process challenges
  • Used control m to develop jobs to schedule the jobs
  • Build the Logical and Physical data model for snowflake as per the changes required
  • Define roles, privileges required to access different database objects.
  • Worked on various kinds of transformations like Expression, Aggregator, Stored Procedure, Java, Lookup,Filter, Joiner, Rank, Router, and Update Strategy. Developed reusable Mapplets and Transformations.
  • Used debugger to debug mappings to gain troubleshooting information about data and error conditions.
  • Involved in monitoring the workflows and in optimizing the load times. Used Change Data Capture (CDC) to simplify ETL in data warehouse applications. Involved in writing procedures, functions in PL/SQL
  • Extensively worked on Views, Stored Procedures, Triggers and SQL queries and for loading the data (staging) to enhance and maintain the existing functionality.
  • Done analysis of Source, Requirements, existing OLTP system and identification of required dimensions and facts from the Database.
  • Created Data acquisition and Interface System Design Document. O Designed the Dimensional Model ofthe Data Warehouse Confirmation of source data layouts and needs.
  • Deploy various reports on SQL Server 2005 Reporting Server O Installing and Configuring SQL Server 2005 on Virtual Machines
  • Used SNOW PIPE for continuous data ingestion from the $3 bucket. Developed snowflake procedures for executing branching and looping Created clone objects to maintain zero-copy cloning. performed Data validations have been done through information schema. Performed data quality issueanalysis using Snow SQL by building analytical warehouses on Snowflake Experience with AWS cloud services: EC2, $3, EMR, RDS, Athena, and Glue Cloned Production data for code modifications and testing
  • Data pipeline consists of Spark, Hive and Sqoop and custom build Input Adapters to ingest, transform and analyze user behaviour (clickstream) data.
  • Good experience in designing and developing CICD pipelines through Azure DevOps.Azure services: Azure HDInsight, Azure Databricks, Azure Data Factory & Azure SQL DW
  • Design and Developed data transformations Pipelines in AZURE Data Factory reverse engineer through existing Talend ETLs.
  • Worked with teams in Data Modelling of new data warehouse Azure Data Lake.
  • Develop Measures and KPI in deploy Azure Analysis Service. Refresh Azure SSAS Models through REST API call using MuleSoft
  • Architect & implement medium to large scale BI solutions on Azure using Azure Data Platform services (Azure Data Lake, Data Factory, Data Lake Analytics, Stream Analytics, Azure SQL DW, HDInsight/Databricks, NoSQL DB).
  • Exploring with the Spark improving the performance and optimization of the existing algorithms in Hadoop using Spark Context, Spark-SQL, Data Frame, and Pair RDD's.
  • Worked on Deep Neural Networks like ANN along with Natural language processing algorithms for Text mining processes
  • Implemented Random Cut Forest, Isolation Forest and Deep Auto Encoder Models for Anomaly detection in both batch and streaming data capture outliers and improve overall quality of strategic data products
  • Applying Hadoop map reduce performance tuning techniques and build hive queries efficiently
  • Designing the distribution strategy for tables in Azure SQL data warehouse
  • Worked on integrating Angular based application to database using NodeJS.
  • Performed Map Reduce Programs those are running on the cluster.
  • Developed multiple MapReduce jobs in java for data cleaning and pre-processing.
  • Written programs in Python for creating the External Tables in the Glue for the respective tables which are located in the S3 buckets to use it in theAmazon Spectrum.
  • Written programs in Spark usingPython, PySparkandPandaspackages for performance tuning, optimization and data quality validations.
  • Worked on developingKafka ProducersandKafka Consumersfor streaming millions of events per second on streaming data.
  • Implemented a distributing messaging queue to integrate withCassandrausingApache Kafka.
  • Hands on experience on fetching the live stream data fromUDBinto HBase table usingPySpark streamingandApache Kafka.
  • Analyzed large and critical datasets usingHDFS, HBase, Hive, HQL, Pig, SqoopandZookeeper.
  • Involved in loading data from UNIX file system to HDFS using Shell Scripting.
  • Used Elasticsearch for indexing/full text searching.

Environment: SQL, Hadoop, Docker, NodeJS,HDFS, Python, SQL, Web Services, MapReduce, Spark, Kafka, Hive, Yarn, Pig, Flume, Zookeeper, Sqoop, UDB, Power BI, Tableau, AWS, GitHub, Shell Scripting.

Confidential - Houston, TX

Data Engineer/data modler

Responsibilities:

  • Used Python 3.X (NumPy, SciPy, pandas, scikit-learn, seaborn) and Spark 2.0 (PySpark, MLlib) to develop variety of models and algorithms for analytic purposes.
  • Design and code required Database structures and components
  • Build the Logical and Physical data model for snowflake as per the changes required
  • Experience on working various distributions of Hadoop like CloudEra, HortonWorks and MapR.
  • Conducted one-on-one sessions with business users to gather data warehouse requirements
  • Analyzed database requirements in detail with the project stakeholders by conducting Joint Requirements D
  • Developed a Conceptual model using Erwin based on requirements analysis
  • Developed normalized Logical and Physical database models to design OLTP system for insurance applications
  • Created dimensional model for the reporting system by identifying required dimensions and facts using Erwin r7.1
  • Identified, formulated and documented detailed business rules and Use Cases based on requirements analysis
  • Exhaustively collected business and technical metadata and maintained naming standards
  • Used Informatica Designer, Workflow Manager and Repository Manager to create source and target definition, design mappings, create repositories and establish users, groups
  • Worked on the snowflaking to reduce redundancy
  • Developed Data Mapping, Data Governance, Transformation and Cleansing rules for the Master Data Management Architecture involving OLTP, ODS and OLAP
  • Developed and maintained data dictionary to create metadata reports for technical and business purpose.
  • Worked at conceptual/logical/physical data model level using Erwin according to requirements.
  • Enforced referential integrity in the OLTP data model for consistent relationship between tables and efficient database design
  • Performed Data Cleaning, features scaling, features engineering using pandas and NumPy packages in python and build models using deep learning frameworks.
  • Involved in all the steps and scope of the project reference data approach to MDM, have created a Data Dictionary and Mapping from Sources to the Target in MDM Data Model.
  • Developed Automation Regressing Scripts for validation of ETL process between multiple databases like AWS Redshift, Mongo DB, T-SQL, and SQL Server using Python.
  • Created PySpark code that uses Spark SQL to generate dataframes from avro formatted raw layer and writes them to data service layer internal tables as orc format.
  • Used the derivatives tool to forward and swaps method with some clients in some assets and also analysing the stock price of different companies and investing in them .
  • Involved in Normalization and De-Normalization of existing tables for faster query retrieval.
  • Developed python programs and excel functions using VB Script to move data and transform data.
  • Live locator and tracker of supply chain. Automated updates when materials running low and need restocking.
  • Developed spark scripts, UDFs using both data frames/SQL and RDD/MapReduce in Spark for data aggregation, manipulation, and ordering and finally put that back into OLTP through Sqoop.
  • Extract Real time feed using Kafka and Spark Streaming and convert it to RDD and process data in the form of Data Frame and save the data as Parquet format in HDFS.
  • Automated the process to copy files in Hadoop system for testing purpose at regular intervals.
  • Coordinate with all the stakeholders and get sign offs, development of data Mapping Documentation
  • Designed data models and data bases and workflows for handling Big data volume such as maintaining Trust Accounting Data.
  • Experience providing technical leadership and mentoring other engineers for best practices on requirements for project.
  • Worked with ETL tools Including Talend Data Integration, Talend Big Data, Pentaho Data Integration and Informatica.
  • Setting up of data models and creating actual data lake onAthena from S3 for visualisation in aws quicksight.
  • Updated Python scripts to match training data with our database stored in AWS Cloud Search, so that we would be able to assign each document a response label for further classification.
  • UsedKafkaproducer to ingest the raw data intoKafkatopics run theSpark Streamingapp to process clickstream events.
  • Performed data analysis and predictivedata modelling.

Environment: AWS, Python, Pyspark, Frameworks, SQL, Spark,Talend, Kafka, Spark Streaming,AWS Redshift, Oracle 10g, Mongo DB, T-SQL, SQL Server

Confidential

Data Engineer

Responsibilities:

  • Used AnyPoint Runtime Manager for deployment and scheduling. Developed reusable flow for complex data processing
  • Configured Hadoop tools like Hive, Pig, Zookeeper, Flume, Impala and Sqoop.
  • Developed an automated process in python to create Unix Shell scripts to perform Hadoop ETL functions like Sqoop, create external/internal Hive tables, initiate HQL scripts etc.
  • Designed data models and data bases and workflows for handling Big data volume such as maintaining Trust Accounting Data.
  • Experience providing technical leadership and mentoring other engineers for best practices on requirements for project.
  • Transforming all reports, dashboards previously built in Excel, SSRS into Tableau.
  • Wrote complex SQL queries using inner join, left join and temp table to retrieve data from the database for reporting purpose.
  • Worked on migrating MapReduce programs into Spark transformations using Scala.
  • Designed, developed data integration programs in a Hadoop environment with NoSQL data store Cassandra for data access and analysis.
  • Resolving issues related to Enterprise data warehouse (EDW), stored procedures in OLTP system and analyzed, design and develop ETL strategies.
  • Pythonwas used in automation of Hive and Reading Configuration files.
  • Involved in Spark for fast processing of data. Used both Spark Shell and Spark Standalone cluster.
  • Processed the image data through the Hadoop distributed system by using Map and Reducethen stored into HDFS.

Environment: Python, Hive, Hadoop, MapReduce, SQL,SSRS, MS Office, MS Excel, Tableau, MY SQL, NO SQL- Cassandra

Confidential

Data Engineer

Responsibilities:

  • Good knowledge of troubleshooting and tuningSparkapplications andHivescripts to achieve optimal performance.
  • UsedSpark Data FramesAPI over Cloudera platforms to perform analytics onHivedata and used Spark Data Frame operations to perform required validations in the data.
  • Performed tuning of SQL queries and Stored Procedure for speedy extraction of data to resolve and troubleshoot issues in OLTP environment.
  • Created several applications for the purpose of Statistical Modelling and Data mining using SAS/Base, SAS/SQL, SAS/Stat, SAS/Graph and also automated applications using SAS/Macros.
  • Involved in administration of data warehouse using Warehouse Administrator functionality of SAS
  • Design and development of different data models according to user specifications in the development of databases for small applications.
  • Created and implemented detailed Staffing Model for Security Solutions Business to evaluate and optimize internal and subcontractor staffing requirements based on the observed and forecasted backlogs, pipeline and historical demand analysis using time series forecasting, and optimization models.
  • Worked with Services and Portal teams on various occasion for data issues in OLTP system
  • Designed and implemented time series forecasting models for estimating server provisioning on the cloud environment which resulted in accurate demand forecasting for cost optimization and budgeting.
  • Environment:Spark, Hive, SAS,Data Warehouse, SQL, OLTP, MS Office, MS Excel, Windows

We'd love your feedback!