PySpark Developer
This is a remote position.
Please go through the entire job post thoroughly before pressing Apply. Post pressing Apply, you shall reach the assessment page that must be attempted.
. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
Busigence is a Decision Intelligence Company. We create decision intelligence products for real people by combining data, technology, business, and behaviour enabling strengthened decisions.
PySpark Developer
Team: Engineering
Location: Remote
Relevant Exp: 0-4 Years
Background: Been there-Done that
Compensation: Above industry standards
Requirements
This is an immediate requirement. We shall have an accelerated interview process for fast closure - you would required to be proactive and responsive Remote position (work-from-anywhere)
Immediate joiners must apply
Data Engineering Experienced - course/competitions/internships/job ( Competitive compensation
************
| MUST HAVE |
************ 1. Code in Python3 - Numpy?
2.Code in Python3 - Pandas?
3.Code in PySpark3 - Core?
4. Code in PySpark3 - SQL?
5.Developed data engineering pipelines on real-world problem (not just toy projects)?
6.Implemented advanced SQL queries
7.Developed complex logics in PySpark3
8.Confidence to learn PySpark3 -MLlib within two weeks?(we shall guide but won't spoon-feed)
===========================================
We are offering one of the most challenging & exciting work on Data Pipelines and Machine Learning Pipelines. You shall be working on sophisticated platforms, products and applications
===========================================
We are looking for developer with real passion for data science pipelines. This is a specialist and individual contributor role. Product development experience preferably at a startup or a lean team is desired
ROLE
We are looking for engineers with real passion for distributed computing with actual hands-on experience developing data application on PySpark. You would be required to work with our data science team on development of several data applications.
Mandatory
1. Must be able to fetching data from data sources (databases, APIs, flat files, etc.)
2. Must know in-and-out of functional programming in Python with strong flair for data structures, linear algebra, & algorithms implementation
3. Must be able to convert, break, & distribute existing Python codes to functional programming syntax 4. Must have worked on atleast one real world project in production on PySpark 5. Must have implemented complex mathematical logics through PySpark at scale on parallel/distributed clusters 6. Must be able to recognize code that is more parallel, and less memory constrained, and you must show how to apply best practices to avoid runtime issues and performance bottlenecks 7. Must have worked on high degree of performance tuning, optimization, configuration, & scheduling in PySpark 8. Must have integrated APIs, streams, databases, files (JSON, XML, CSV etc) through PySpark
Preferred 1.Good to have working knowledge vinaigretteof first-class, high order, & pure functions, recurisons, lazy evaluations, and immutable data structures
2. A firm understanding of the underlying mathematics will be needed to adapt modelling techniques to fit the problem space with large data (1M+ records)
3. Good to have worked on PySpark MLlib and PySpark ML 4. Configured Checkpointing and Directed Acyclic Graphs (DAG) on PySpark cluster
5. Worked on development of data platform
Benefits
How to Apply You should apply online by clicking "Apply Now". For queries regarding an open position, please write
For more information, visit
Products:
Careers:
Research:
Jobs:
Scaling established startup innovating & disrupting various domains through artificial intelligence. We bring those people onboard who are dedicated to deliver wisdom to humanity by solving the worldβs most pressing problems differen